Skip to content

Zanwen Fu / Engineer & FounderMake something
people want

Agents are easy to demo and hard to depend on. I build the ones people depend on.

I founded VYNN AI (opens in new tab), a personal financial analyst, and built every part of it myself. More than 5,000 investors have signed up. Every number in it is computed, not generated. When Wall Street disagrees, it stands by its own and shows you both.

Before that I was an early employee at AutoCodeRover, the code-repair agent acquired by Sonar (opens in new tab). I’ve also shipped agentic AI at Robinhood and reliability infrastructure for Binance’s Web3 Wallet.

What I took from all of it: the model is rarely what decides whether an agent holds up. Everything around it is. That’s what I build now: Agent OS, an operating system for AI agents, and Errata-Bench, coming soon.

01Work

Selected work

All projects & research

VYNN AI

A personal financial analyst for every investor

Founder & sole engineer / 2025—present

I started VYNN because the research that moves markets is sold by the seat, and the AI tools built to replace it make up their numbers. So I built one that shows its work: name a company, fund, coin or prediction market, and in about two minutes VYNN returns a sourced report, a live Excel model and a dated thesis, for about three cents. More than 5,000 investors have signed up. Next, it will watch their portfolios around the clock.

Engineering decision

The language model never writes a number. Code computes every figure, and a report that cites less than 95% of its claims gets sent back. When VYNN disagrees with Wall Street, it stands by its number, flags the gap and shows both.

Inside VYNN AIArchitecture, financial modeling & production systems
VYNN / Personal financial analyst
Research trajectoryIllustrated execution
Investor

Analyze NVIDIA: valuation, recent news, and key risks.

Research agentSelect tools · read results · continue
01get_financialsNVDA

Statements · cash flow · estimates · market data

02write_reportNVDA
Model generationValuation engine

Formula-backed DCF
Editable workbook

News analysisEvidence pipeline

Screen & summarize
Source-linked findings

Concurrent specialists
Report & publication checksSources Arithmetic Method agreement
↳read_reportReuse saved research

Start with the investor’s question

The research agent selects tools from the request, reads their results, and decides what to do next. This example asks for a full company analysis.

01 / 06Choose tools
Tool orchestration source
Tool selection · parallel research · verified outputSystem details

AutoCodeRover

Autonomous code repair, inside the IDE

Research engineer / 2024—25

In early 2024, before Claude Code or Codex, AutoCodeRover was one of the first agents that could take a real GitHub issue and fix it on its own. I joined early, worked on lifting it to 51.6% on SWE-bench Verified, built its Self-Fix loop, and built the JetBrains plugin that put it inside the IDE. Sonar acquired it in February 2025.

Engineering decision

Developers don’t stop typing while an agent works. So the plugin merges the agent’s fix into their latest code on the syntax tree, keeping every edit they made in the meantime.

Inside AutoCodeRoverThe plugin, repair workflow & research
AutoCodeRover / JetBrains plugin
GumTreeThree-way AST merge
SimpleNamelookup → findKeep local
InfixExpression&& → ||Take patch
Baseline codeRepository.java
User lookup(String key) {
  if (key == null && key.isEmpty())
    return User.missing();
  return cache.get(key);
}

Three versions become three syntax trees

The plugin parses the Git baseline, your working copy, and the repair applied to that baseline with GumTree’s JDT parser. The tree excerpt below follows one method.

01 / 04Parse
GumTree mapping & merge implementation
JDT parsing · GumTree mappings · merged ASTInside the plugin

Errata-Bench

Coming soon

Agent OS

An operating system for AI agents, with git as its memory

Creator / open source / 2026

Every team building agents rebuilds the same plumbing: memory, rollback, checks, budgets. Agent OS is the operating system I’m building underneath them. You hand it a goal; a central brain plans, each piece of work runs in its own process, and a monitor that can’t touch anything signs off. Every step is a commit in one git repository, so nothing unverified ships and nothing is lost. It already runs real benchmark tasks end to end: 29 trials on 16 coding tasks, every one accepted by the benchmark’s admission check.

Engineering decision

Whoever does the work doesn’t grade it. A worker’s word isn’t evidence; only a monitor’s sign-off on the exact end state is, and only signed-off work is delivered. A failed attempt is rolled back but never erased, so the next plan can learn from it.

  1. aControlthe central brain

  2. bExecutionone process per contract

  3. cMemoryone git repository

Follow a goal through the system

You give it a goal. A central brain breaks the goal into contracts, a separate worker process carries out each one, and an observe-only monitor certifies what it finished. Every plan, model call, command, verdict and rollback is a commit in one git repository. Work that was not certified is never delivered, and nothing that happened is lost.

Point at any part, or play the walkthrough, to see what it does, what it leaves on record, and what has been shown of it so far.

The repository’s rollback demo, which CI runs on every push

EVAL id=step-02 passed=False reason=`pytest -q` exited 1  sha=4e88e78
REV  id=step-02 to=d59174c remaining_retries=2
EVAL id=step-02 passed=True  reason=`pytest -q` exited 0  sha=d407440

Caught and rolled back. A scripted worker breaks the tests on its second step; the retry lands clean. From the single-process kernel this runtime grew out of.

10 parts

The system, following a goal. The central brain plans and decides (a). Each contract runs in a worker process of its own, watched by a monitor that cannot act (b). What they do becomes commits in one git repository (c): a checkpoint, certified, a failed attempt.

Every part, as text
  1. Goal (criteria, limits). A task in plain words, the criteria that must hold at the end, a budget and a deadline.
    1. success criteria: What must hold at the end. The planner has to account for every one, with evidence.
    2. limits: The budget, the deadline, the models and the call limits are fixed on admission. A plan cannot raise them.
    3. on record: The goal is recorded before anything else happens.
    Shown so far: In the real-model trials each goal was a benchmark task, run as published in its own container.
  2. Planner (one call per plan). One model call per plan. It is given the goal, its criteria and the world on record: every branch, and the outcome of every worker so far.
    1. one contract per worker: The task, the outputs to produce, and the criteria its monitor will judge.
    2. judges every criterion: Each plan assesses every goal criterion, with the evidence for it.
    3. reads only the record: It never sees a worker’s private context, only what is on record.
    4. on record: Each plan is a commit on control. A plan that fails to parse changes nothing.
    Shown so far: In the first real-model trials, a completed task took two plans.
  3. Runtime (the durable cycle). The loop that makes plans happen. It waits while workers work, and asks the planner to revise when something gives it a reason to: a finished worker, a failure, a refusal.
    1. collect and deliver: It collects the report of each worker that ended and delivers certified outputs into integration.
    2. replan or finish: A failure, a refusal or a stopped worker is a reason to replan, not a reason to stop.
    3. bounds: A run ends when the plan is complete, or at the clock, the budget, or a limit on plans.
    4. on record: Each effect is recorded before it happens and after. After a crash it is looked up, not repeated.
    Shown so far: In a drill where working time ran out, the run still closed with a reply that named what was run and what was not verified.
  4. Supervisor (owns the processes). It owns the worker processes: it prepares a branch for each contract, starts the worker in its own process group, enforces the deadline, and reaps the whole process tree.
    1. start: One operating-system process per contract, on its own branch of memory.
    2. stop: A worker is asked to stop before it is killed, so it can settle a paid call and write its report.
    3. reap: The whole process tree is reaped at the end. No process escapes.
    4. on record: A start is recorded before the process exists; after a crash the process is found by its identity.
    Shown so far: When the benchmark’s own time limit cancelled the agent, the trial was sealed within 5 seconds, with its commands recorded in order.
  5. Worker (its own process). An operating-system process with one contract, one branch of memory and a set of tools.
    1. contract in: It starts from integration, with the task, the outputs to produce and the criteria.
    2. the loop: Model call, tool call, result, again.
    3. claim out: It ends with a structured claim: completed, blocked, or needing another turn, with its evidence.
    4. on record: Checkpoints on its own branch: the files, the transcript so far, and the reason.
    Shown so far: On a benchmark task, one worker at a time works in the task’s container.
  6. Terminal broker (task environment). When workers act on a container they share, every command goes through a broker. Workers share the environment; the brain stays out.
    1. one at a time: One command runs at a time, whichever worker sent it.
    2. timeouts: A command past its timeout is killed with the processes it started. Its output is kept and the container stays usable.
    3. uncertain: A command whose outcome cannot be confirmed is recorded as uncertain and never run again on a guess.
    4. on record: Each command, its exit status and its output enter a ledger before the worker sees the result.
    Shown so far: In a drill, a command that ignored termination and left a detached child was ended alone in 1.1 seconds; the container stayed usable.
  7. Model API (priced, budgeted). Every model call, from the planner, the workers and the monitors, goes through one facade.
    1. priced: A model with no verified price cannot be called.
    2. budgeted: Before a call is sent, its worst-case cost is reserved against what is left of the budget.
    3. paid once: After a crash the stored reply is replayed; the provider is not asked, and not paid, a second time.
    4. on record: The reply is journaled before anything acts on it.
    Shown so far: The trials ran gpt-6-astra on Azure OpenAI in every role. A run stopped when its working time ran out still reported its exact cost.
  8. Monitor (observe-only). Every worker has a monitor that only observes. It has no tools, so it cannot fix what it finds: it can only say so.
    1. reads the transcript: As it grows, not only at the end.
    2. sends verdicts: The worker is shown them and cannot claim completion until it has answered.
    3. certifies, or refuses: It judges the exact end state against the contract. Only certified work can be delivered.
    4. on record: A verdict is a note attached to the exact state it judges.
    Shown so far: With the evidence itself in front of it, the certifier finished a long task in under five minutes for about $4.40, against 17 minutes and $18.58.
  9. Memory (one git repository). A session’s memory is one git repository. What other harnesses add as side files is already there, with history, diffs and merges.
    1. a branch per actor: control holds the plans, integration only certified results, and each worker has its own.
    2. a commit per checkpoint: Files, the transcript so far, and the reason.
    3. rollback is an append: A new commit restores an earlier state. The failed attempt stays readable for the next plan.
    4. typed merges: Handing work over is a typed merge. A conflict is a value to act on; no model invents a resolution.
    Shown so far: The memory layer has its own test suite, with a test for every defect an adversarial review found, and a crash-consistency demo.
  10. Result (reply + record). The final reply, the certified outputs and the complete record.
    1. the final reply: Written by the planner when every criterion is met.
    2. a closing reply: If time or budget runs out first, it says what was done and what was not.
    3. nothing unverified: Only work a monitor certified reaches integration.
    4. on record: Nothing is deleted: every plan, call, command, verdict and rollback stays in the repository.
    Shown so far: 29 real-model trials of 16 coding tasks, all accepted by that benchmark’s own admission check.

LUMINA

Automating the manual screening of medical reviews

First author / research / 2024—25

A medical systematic review takes a team about 15 months, starting with thousands of abstracts read by hand to find the few that belong. I built LUMINA to do that screening: four agents that sort, check and challenge each other’s calls. Across 15 published reviews it missed 8 of 378 included studies and cut full-text reading to under 9%. I’m the first author.

Engineering decision

A missed study changes the evidence; an extra read costs ten minutes. So the first pass keeps anything uncertain, and a reviewer sends doubtful calls back for revision, up to three rounds.

Inside LUMINAAgent architecture, methods & evaluation
LUMINA
Two-tier screeningCitation lifecycle
Title & abstractCitation A

Illustrated inclusion path

trace.jsonlDecision · feedback · revisions · cost

Begin with the title and abstract

A citation enters alongside the systematic review’s scope. The first tier asks whether it could be relevant, keeping uncertain studies in play.

01 / 08Read citation
Screening pipeline source
Four agent roles · two screening tiersFull architecture
All projects & research Architecture, experiments, and source code

02Experience

Experience

Robinhood

· Agentic AI Team

May 2026 – Aug 2026

Machine Learning Engineer Intern · Menlo Park, CA

Worked on Robinhood Cortex, the AI assistant for Robinhood Gold, with a focus on adoption and engagement. Solo-designed and shipped a proactive agent that decides when something deserves a customer’s attention instead of waiting to be asked. Cut false positives five-fold in backtesting with no material event missed, scaled coverage 30× at flat latency, and caught and fixed a production delivery failure that monitoring had missed.

Python · Go · Kafka · Kubernetes · Airflow · DynamoDB

Jul 2025 – present

Founder & Sole Engineer · Durham, NC

Founded VYNN to give every investor a personal financial analyst, and built every part of it: the research agents, the valuation engine, the backend, the app and the infrastructure it runs on. More than 5,000 investors have signed up. It turns six to twelve hours of an analyst’s research into a sourced report and a live Excel model in about two minutes, for about three cents. Code computes every number, and every release is checked against a set of reference valuations before it ships.

Python · LangGraph · FastAPI · React · TypeScript · MongoDB · Redis · Docker · Hetzner

Research Engineer · Singapore

Worked on lifting the agent to 51.6% on SWE-bench Verified, building Self-Fix and targeted replay into the repair backend. Built the JetBrains IDE plugin end to end: a three-way merge on the syntax tree, SonarLint analysis sent straight to the agent, and live streaming with per-step feedback. Sonar acquired AutoCodeRover in February 2025; the same team’s Foundation Agent later reached the top of the unfiltered SWE-bench leaderboard at 79.2%.

Kotlin · Python · IntelliJ Platform SDK · GumTree · SonarLint · JGit · OkHttp

Binance

· Web3 Wallet Team

Jul 2025 – Oct 2025

Software Engineer Intern · Singapore

Built API automation and performance infrastructure for the Web3 Wallet team. Integrated regression checks into CI/CD, engineered concurrent workloads to exercise backend services under load, and instrumented monitoring to surface data-consistency failures before release.

Java · REST APIs · CI/CD · JMeter · Observability

Aug 2025 – Apr 2026

Graduate Teaching Assistant · Durham, NC

Designed and ran CS 590 (Software Development Studio), where graduate students build AI debugging agents inspired by AutoCodeRover and deploy full-stack applications. Mentored teams in CS 408 and CS 390 on software architecture, DevOps, and LLM-oriented programming, shipping production software for outside clients.

Python · Docker · CI/CD · Git/GitLab · Web Assembly

Jan 2024 – Jul 2025

AI Researcher · Singapore

First author of LUMINA, a four-agent system that screens medical literature for systematic reviews; designed and built it and ran the evaluation. Across 15 published reviews it missed 8 of 378 included studies (98.2% mean sensitivity, 87.9% specificity), a 35× lower miss rate than a reimplementation of Li et al. (2024), at under a cent per citation. The pipeline has screened about 150,000 candidate abstracts in all.

Python · gpt-4o-mini · o3-mini · PICOS · Evaluation

Earlier

ST Engineering

May 2023 – Aug 2023

Full-Stack Software Engineer

Built a full-stack railway dashboard connecting backend services to live train-status reporting across Singapore’s MRT network.

Quantum Software Engineer

Built a quantum experiment compiler and controller for FPGA/DDS hardware, with threaded execution and live data visualization.

NUS Computing

Feb 2024 – Nov 2024

Research Assistant

Built a dynamic web scraper and developed a chatbot for a research project at NUS Computing.

03About

Research, industry & ownership

I like taking a product from the first useful version through to the details that make it dependable. With VYNN (opens in new tab), that has meant everything from financial models and worker orchestration to the way a missing number appears on screen.

My work has covered code repair at AutoCodeRover, research screening with LUMINA (opens in new tab), proactive AI at Robinhood, and backend reliability engineering at Binance. I’m now building Agent OS (opens in new tab), an operating system for AI agents that keeps everything they do on record and delivers only what has been verified. Errata-Bench is coming soon.

I studied computer science at NUS and am pursuing my master’s at Duke, where I’ve also taught software engineering.

Duke University

M.S. Computer Science · AI/ML
Scholar profile (opens in new tab)

2025–2027

National University of Singapore

B.Comp. Computer Science · Honours, Distinction
Distinction in Software Engineering (opens in new tab) · HKU exchange

2021–2025

Open to full-time engineering roles from May 2027.

05Teaching

Teaching software engineering

Graduate TA at Duke CS. Across three courses I've run the hands-on infrastructure side: 48 students across 15 teams shipping systems with Docker, CI/CD, and AI agents, on labs and benchmarks I build and maintain.

Graduate course covering production software engineering, Docker, CI/CD, API design, and server deployment — culminating in students building an AI debugging agent (inspired by AutoCodeRover) and a full-stack social media application.

Mentoring teams on architecture, testing, DevOps, and full-stack development to deliver production-ready software for outside clients.

Leading weekly labs covering AI agents, LLM-oriented programming, Docker, APIs, and system design.

Teaching materials — labs, benchmarks, and the LLM-teammate pipeline →