Skip to content
Back to selected work

Projects & research

Products I’ve built, systems I’ve designed, and experiments I’ve run. Each project includes its architecture, results, and source code.

Selected work

each with a page of its own

More research

security, ML systems, fine-tuning, evaluation

Four shorter studies, each with its code and data. They ask what stops a prompt injection in a deployed agent, why speculative decoding fails on a small GPU, how much of a LoRA adapter is really used, and whether a fine-tuned model predicts football or remembers it.

architectural-damping

Security study · Duke, spring 2026

Do the deterministic parts of an agent stop prompt injections that fool its language model? In VYNN’s news pipeline, every poisoned article changed the model’s screening output, but only 2 of 12 changed the rating a user would see: the calculator after the model absorbed the rest. Reading the calculator’s source predicted which cases could break (6 of 6 held) and bounds one poisoned document’s effect at 7.3 points. A defense that separates instructions from data stopped the remaining attacks in the test (0 of 9), but an adaptive attacker found a new route.

The setup

VYNN’s news path, run against an offline copy of its database: a gpt-4o-mini screener reads articles and feeds a deterministic calculator that sets the rating band. The benchmark has 20 clean and 60 poisoned cases, one poisoned document each. No poisoned content ever reached real users.

What the calculator absorbs

Every poisoned case changed the screener’s structured output. Only 2 of 12 pilot cases changed the rating, a damping ratio of 0.83, and direct-override attacks changed none. From the calculator’s weights, one poisoned document can move the score by at most 7.3 points; the largest shift observed was 1.9.

Defenses, and an adaptive attacker

Separating instructions from data (struq-lite) cut attack success from 22% to zero on a matched 9-case slice. A second model asked to verify the screener caught nothing. In one adaptive round the attacker kept the same success rate by switching attack families, and the defense let through an attack the undefended pipeline had resisted.

Report and artifacts

The technical report (opens in new tab) covers the threat model, the attack taxonomy and every case. Frozen manifests and a deterministic model cache let anyone recheck the numbers.

Python · LangGraph · gpt-4o-mini · Claude Sonnet 4 · MongoDB · pytest

zanwenfu/architectural-damping

speculative-decoding-t4

ML systems study · Duke, spring 2026

On an NVIDIA T4, speculative decoding with a draft model gained nothing over plain decoding: Sequoia’s cost model predicted 1.68×, and I measured 0.56×. Keeping the cache between steps, the obvious fix, made it worse (0.46×). One number explains it: on this bandwidth-bound GPU a 1B draft model costs 60% of a 3B model per call, not the 33% its size suggests. A cost model calibrated on that matches the measurement within 1.1%. Only prompt lookup, which needs no draft model, sped decoding up (1.28–1.39×).

What was measured

Llama-3.2 1B drafting for 3B in fp16 on one T4: linear speculative decoding, the Sequoia tree with and without a persistent cache, and prompt lookup, over 10 prompts and 30 trials each. Every method’s output was checked token for token against plain decoding: 20 of 20 identical.

Why the cost model was wrong

Sequoia’s cost model assumes a draft call costs what the draft’s size suggests. On the bandwidth-bound T4 every call pays a fixed cost to load weights: a draft call took 20.0 ± 0.2 ms whether its cache was short or 8.5× longer. That makes the 1B draft cost 60% of the 3B target per call, not 33%. With that ratio, the calibrated model predicts 0.563× against 0.557× measured.

A rule you can use

The repository ships the calibrated cost model as a small Python advisor: give it a draft model, a target and a GPU, and it predicts whether speculative decoding will help before you spend the compute.

Python · PyTorch 2.10 · Transformers · CUDA · Jupyter

zanwenfu/speculative-decoding-t4

svd-lora

Fine-tuning study · Duke, autumn 2025

Trained LoRA adapters use far less rank than they are given. Cutting each adapter to the rank that keeps 90% of its energy, then training briefly, keeps SST-2 accuracy and raises IMDB F1 from 0.872 to 0.892, with about 30% of the parameters.

Method

Train standard LoRA at rank 8 on DistilBERT. Take each adapter’s singular values and keep the smallest rank that holds 90% of its energy, which by the Eckart–Young–Mirsky theorem is the best approximation at that rank. Then train two more epochs in the smaller space. The compression step is about 30 lines.

Results

On SST-2, accuracy 0.893 against LoRA’s 0.892. On IMDB, F1 0.892 against 0.872 and accuracy 0.881 against 0.868. The average rank was 2.4, selected on SST-2: about 30% of LoRA’s parameters.

Against newer variants

In a comparison with HiRA and Sparse LoRA, neither variant beat plain LoRA on accuracy, and neither ran faster.

Availability

Open source (MIT). Runs on CUDA, Apple MPS or CPU, with PEFT and Hugging Face.

Python · PyTorch · PEFT · Hugging Face · DistilBERT · SVD

zanwenfu/svd-lora

football-llm

Applied ML audit · 2026

Can a fine-tuned LLM predict World Cup matches, or does it just remember them? My spring model, Llama 3.1 8B fine-tuned on player statistics, seemed to beat XGBoost by more than 20 points. Then I hid the team names: exact-score accuracy fell from 43.8% to 10.9%, and 59% of its predictions were the real score or its mirror. It had memorized the 2022 tournament, which its pretraining data covers.

The model

A QLoRA adapter on Llama 3.1 8B reads player statistics for both starting line-ups and predicts the final score, before kickoff or at halftime. It was trained on the 2010, 2014 and 2018 World Cups and tested on the 64 matches of 2022.

The audit

Three tests say the headline was mostly memory. With the names hidden, exact scores fall from 43.8% to 10.9% (paired McNemar p < 0.001) while winner accuracy does not change. With names, 59% of predictions are the real score or its home-away mirror, against about 10% by chance. And the backtest rewards its own assumptions: a rule with no model returns +1,125% under the same simulation.

What survives

The halftime score is real information: with names hidden, over/under accuracy rises from 56% to 70% at halftime. A one-line rule captures it at least as well.

Companion study: coherent answers only

A sister repository trains adapters on 48, 96 and 192 examples and sorts every generation into six kinds, from coherent success to unparseable. Lenient parsing flatters fine-tuning: QLoRA scores 61.7% against 49.2% for 5-shot prompting, but counting only coherent answers, both score 42.2%.

Reproduce

Every number in the audit reproduces in about five seconds, without a GPU.

Python · QLoRA · Llama 3.1 8B · XGBoost · McNemar · Hugging Face

zanwenfu/football-llm

Tooling, prototypes, course artifacts, and in-progress work all live at github.com/zanwenfu (opens in new tab) — the curated story is above, the full archive is a click away.