Projects & research
Products I’ve built, systems I’ve designed, and experiments I’ve run. Each project includes its architecture, results, and source code.
Selected work
each with a page of its own- VYNN AIFounder & sole engineer · 2025—presentA personal financial analyst for every investor: a sourced report, a live Excel model and a dated thesis in about two minutes, for about three cents. I founded it and built every part of it. More than 5,000 investors have signed up.
- AutoCodeRoverResearch engineer · 2024—25 · acquired by SonarOne of the first coding agents, fixing real GitHub issues in 2024. I worked on lifting it to 51.6% on SWE-bench Verified and built its IDE plugin; Sonar acquired it in 2025.
- Errata-BenchComing soon
- Agent OSCreator · open source · 2026An operating system for AI agents. A central brain plans the work, each piece runs in its own process, and a monitor signs off before anything ships. Every step is a commit in git.
- LUMINAFirst author · research · 2024—25A medical review takes a team about 15 months, much of it screening abstracts by hand. I built four agents that do it: across 15 reviews they missed 8 of 378 included studies and cut full-text reading to under 9%.
More research
security, ML systems, fine-tuning, evaluationFour shorter studies, each with its code and data. They ask what stops a prompt injection in a deployed agent, why speculative decoding fails on a small GPU, how much of a LoRA adapter is really used, and whether a fine-tuned model predicts football or remembers it.
architectural-damping
Security study · Duke, spring 2026Do the deterministic parts of an agent stop prompt injections that fool its language model? In VYNN’s news pipeline, every poisoned article changed the model’s screening output, but only 2 of 12 changed the rating a user would see: the calculator after the model absorbed the rest. Reading the calculator’s source predicted which cases could break (6 of 6 held) and bounds one poisoned document’s effect at 7.3 points. A defense that separates instructions from data stopped the remaining attacks in the test (0 of 9), but an adaptive attacker found a new route.
The setup
What the calculator absorbs
Defenses, and an adaptive attacker
Report and artifacts
Python · LangGraph · gpt-4o-mini · Claude Sonnet 4 · MongoDB · pytest
speculative-decoding-t4
ML systems study · Duke, spring 2026On an NVIDIA T4, speculative decoding with a draft model gained nothing over plain decoding: Sequoia’s cost model predicted 1.68×, and I measured 0.56×. Keeping the cache between steps, the obvious fix, made it worse (0.46×). One number explains it: on this bandwidth-bound GPU a 1B draft model costs 60% of a 3B model per call, not the 33% its size suggests. A cost model calibrated on that matches the measurement within 1.1%. Only prompt lookup, which needs no draft model, sped decoding up (1.28–1.39×).
What was measured
Why the cost model was wrong
A rule you can use
Python · PyTorch 2.10 · Transformers · CUDA · Jupyter
svd-lora
Fine-tuning study · Duke, autumn 2025Trained LoRA adapters use far less rank than they are given. Cutting each adapter to the rank that keeps 90% of its energy, then training briefly, keeps SST-2 accuracy and raises IMDB F1 from 0.872 to 0.892, with about 30% of the parameters.
Method
Results
Against newer variants
Availability
Python · PyTorch · PEFT · Hugging Face · DistilBERT · SVD
football-llm
Applied ML audit · 2026Can a fine-tuned LLM predict World Cup matches, or does it just remember them? My spring model, Llama 3.1 8B fine-tuned on player statistics, seemed to beat XGBoost by more than 20 points. Then I hid the team names: exact-score accuracy fell from 43.8% to 10.9%, and 59% of its predictions were the real score or its mirror. It had memorized the 2022 tournament, which its pretraining data covers.
The model
The audit
What survives
Companion study: coherent answers only
Reproduce
Python · QLoRA · Llama 3.1 8B · XGBoost · McNemar · Hugging Face
Tooling, prototypes, course artifacts, and in-progress work all live at github.com/zanwenfu (opens in new tab) — the curated story is above, the full archive is a click away.