VYNN AI · Agentic AI · LLM Reliability · Valuation · Software ArchitectureCite
VYNN AI, Six Months Later: A Simpler Agent and Stricter Numbers
Since March I swapped VYNN's agent graph for one tool loop, made every number checkable, fixed a measured bias and put two checks before every release.
In March I wrote about building VYNN AI alone and the decisions I got wrong. Six months later the agent is simpler and everything around it is stricter. Most of what made VYNN more reliable happened outside the model. I measured the engine as a whole, checked every figure before it reached anyone, and stopped releases that changed numbers without a reason.
VYNN is a personal financial analyst for every investor. You ask about a company, a fund, a coin or a prediction market, in your own words and in any language. For a company it can run a full analysis, which means a discounted cash flow (DCF) valuation built as an Excel model with live formulas, plus a research report that cites its sources. A DCF values a company by projecting the cash it will generate and discounting that cash back to today. 5K+ people have signed up, and it is free during early access.
March and October, side by side.One loop instead of a graph
In March every request went through a supervisor. It sorted each question into one of four modes, such as a full analysis or a quick news check, and then routed it through a fixed graph of worker agents. That works when people type a ticker. Most people don't. They ask whether Costco is a good buy at today's price, what a rate cut means for their bank stocks, or what a fund actually holds, sometimes in Chinese. Anything outside the four modes got a canned reply.
In August I replaced the supervisor with a single loop. The model reads the question, picks a tool, reads the result, and either calls another tool or answers. The pattern is usually called ReAct, for reason and act, and one turn can take at most eight steps.
The decision that mattered was what became a tool. The four old specialist agents, which gather the financials, build the valuation model, screen the news and write the report, became four of the loop's 18 tools, with their old guarantees intact inside them. The model decides whether a question needs that heavy path. It cannot change how the path computes. The other tools fetch prices, news, macro data, funds, crypto and prediction markets, or price options and portfolio risk.
I wrote the loop by hand, so it keeps the promises the rest of the stack relies on, like streaming progress and ending cleanly when something crashes. The old LangGraph supervisor still sits in the repo behind a switch, but production runs the loop.
The logs show how the loop chooses. Since August, about four in ten questions have run a full analysis, about four in ten were answered with lighter tools, and the rest needed no tool at all. A question that uses tools usually takes one or two rounds of them, and one round can call several tools at once. None has used more than six of the eight steps.
VYNN today. A question enters one agent loop, which runs inside its own container and calls tools. Every number comes from code and published data, and passes checks before it reaches the answer.Speed and cost
Full analyses got faster, though not because of the loop. The 6.4 minutes I quoted in March was a single run from December 2024 on the old pipeline, not an average. The time came back by letting independent work run at once. The news screener sends its batches of articles together instead of one after another, the report's independent sections are written in parallel, and the valuation model is built while the news is analyzed. Since September 27 the median full analysis has taken 1:57 from question to answer, across 34 runs, and the median quick answer 13 seconds.
They got cheaper too. On September 27 the default model became gpt-6-luna. Since then the median full analysis has spent 1.3 cents on model calls, against 6.6 cents in the eight weeks before, and a quick answer about a fifth of a cent.
Every number is computed, then checked
Every number is computed, not generated, as it was in March. What changed is how hard the code checks the model's work before anything reaches you.
A calculator sets the rating and the 3, 6 and 12 month price targets from the valuation. After the model writes the report, a validator compares each rating and target in the text with the calculator's. If the model misstated one, the validator puts the calculator's value back and makes the model rewrite the prose around it. When there is news to cite, at least 95% of the sentences that make a claim must cite a source, and the cited article has to support the sentence. A real citation attached to the wrong claim fails. That failure is common. In one study, about 30% of the statements GPT-4o made with web search weren't supported by the sources it cited.
The run logs kept the validator's verdict for 13 runs, from July and August on an older model. The current report path doesn't record it yet. In none of the 13 did the model misstate a rating or a target. Ten first drafts failed on citations. Six cited too few of their claims, and one or two rewrites fixed all six. Four cited evidence that didn't exist. No rewrite fixed those, so VYNN published a recommendation assembled in code from the evidence instead.
The Excel model has ten sheets, from the raw statements to the valuation summary. In March I said I should replace my homemade formula evaluator with a library. I kept it and made it stricter. Before a workbook is saved, every formula on every sheet is evaluated. A second pass then checks that the numbers agree the way accounting says they must. It exists because a formula can compute perfectly and still point at the wrong cell. One of mine did, and the model's operating profit came out more than $100 billion away from what the company reported. If either pass fails, that number never reaches you as a fair value.
The inputs changed too. Since September the risk-free rate, the government bond yield every valuation starts from, comes from published data, read from FRED for the dollar, the European Central Bank for the euro and Japan's Ministry of Finance for the yen. Country risk comes from Aswath Damodaran's published tables, and every input in the workbook names its source and its date.
How one number travels, using the Costco run of October 2. Inputs come from published sources, code computes the value, the workbook and the validator check it, and the gap to analysts is flagged with both numbers shown.Measuring the engine, not the answers
In September I looked at all the valuations together instead of one at a time. Across the 49 valuations stored by then, the median company came out 15.7% below its market price, 69% came out negative, and the ratings ran 18 sells to 5 buys. Each answer looked defensible on its own. Together they said the engine was miscalibrated.
I found two causes. The first was the equity risk premium, the extra return investors demand for owning stocks instead of government bonds. The workbook printed Damodaran's published premium of 4.23%, but the code discounted at a 5.5% constant I had set long before, and a higher discount rate lowers every DCF at once. The second was the vote. The headline value averaged three methods with equal weight, and two of them were DCFs built on the same inputs. One DCF opinion got two thirds of the vote, and the comparables, which value a company by the prices of similar ones, got a third.
I switched to the published premium and let the two DCFs count as one view, weighed equally against the comparables. On Apple, the weighting fix alone moved the headline value from $168 to $195.
Then I measured again. Across 72 large US companies, the median still came out about 15% below its price. Damodaran derives his premium from today's prices while assuming companies grow at the analysts' five-year rate. VYNN's model assumes slower growth, and that pairing pulls every value down by the same amount. So I calibrated the premium his way on VYNN's own assumptions. It is the premium at which the engine values the median company in a fixed sample of 74 at its market price, and it came out at 3.24%.
That assumes the market prices the median company fairly, as Damodaran's own premium does. It removes the bias every company shares and keeps the differences between companies, which are what VYNN is for.
Old valuations stay dated and unchanged, and those made under the old assumptions are marked in the app instead of quietly rewritten.
When VYNN and Wall Street disagree
A valuation has no answer key. When VYNN's number lands far from what analysts expect, there are two easy ways out. You can drop the number, or you can pull it toward the consensus. VYNN does neither. Since October 2 it flags the gap and shows both.
That day someone asked whether Costco was a good investment at $920.65. VYNN's fair value was $437.75, 52% below the price, while the mean target of 35 analysts was 15% above it. VYNN rated it a strong sell, marked the rating low confidence and stated both numbers. The same alert appears in the chat, on the report's first page and in the workbook. Analyst targets are a benchmark and never enter the valuation math.
When the model can't support a single value, say because a company's latest statements are stale, VYNN answers with the range its methods produce.
Two checks in front of every release
Between September 15 and 26, a change cut how often VYNN put a number on a company. Nothing crashed, so nothing alerted me, and the drop reached users before I caught it. A crash alert can't see that kind of regression, so the check has to look at the outputs.
Now two checks guard every release. Before a deploy, the candidate engine rebuilds the valuations of 16 reference companies, from Nvidia and Microsoft to LVMH and an Indian jeweller, without calling a language model. A file in the repo states what each one should produce and why. A candidate ships only with zero unexplained differences, or with the expectation changed in the same commit as the engine change that explains it. I read that diff before every deploy.
Every night, valuations run again against a golden dataset of 100 companies from the Nasdaq-100, and a release is blocked if any of them drifts beyond its threshold. The production worker image is pinned by its digest, a hash of its contents rather than a tag that can move, so a rollback means pointing back at the previous digest.
How a release is guarded. A change has to explain every difference on 16 reference companies before its image is pinned, and a nightly regression on 100 companies blocks the next release if a value drifts.The decision I kept
In March I called running each analysis in its own container my biggest mistake. I still do it. With limits, the isolation is what keeps one bad run from taking down the rest. Each container gets 768 MB and one CPU, drops every Linux capability and caps how many processes it can start, and at most four run at once.
The limits came from watching the server fail. A run peaks near 350 MB while it pulls five years of history and builds the model. Past what the machine could hold, the kernel killed whichever process it picked, sometimes the API itself, and one traffic spike took VYNN down for everyone. Now a runaway run is killed alone, and an extra request waits for a free slot.
The website
At the end of September I rebuilt vynnai.com around recorded questions. Each example on the homepage is a real question asked in the app, shown with its dated answer and the files behind it. Five companies have a research record built from one run's own report and workbook, and every report and workbook opens in a preview before anything downloads. The site uses the app's typefaces and its paper background, so following a link into the app feels like the same product.
vynnai.com today. The homepage opens with "Know what you own." above NVIDIA's thesis of record as it appears in the app, with its fair value, rating and what the price requires.Where VYNN is today
VYNN covers companies, funds, crypto, prediction markets, options and portfolio risk, across five repositories. The agent, stock-analyst, is open source on GitHub. The API, api-runner, is a FastAPI service with MongoDB and Redis that starts the analysis containers. The web app, vynnai-web, is React and TypeScript, and the website, landing-page, is built with Next.js.
In March I also described a proactive VYNN that would watch the market and tell you when news changes your thesis. That needs a thesis to compare against. Every analysis is now kept as a dated thesis for the person who asked, which is the first piece of it.
You can ask VYNN a question at app.vynnai.com, read its dated research at vynnai.com, or see how it is built at zanwenfu.com/projects/vynn-ai.
Cite this essay
If you build on or quote this essay, please cite it.
Zanwen Fu. 2026. “VYNN AI, Six Months Later: A Simpler Agent and Stricter Numbers.” zanwenfu.com, October 7, 2026. https://zanwenfu.com/blog/vynnai_six_months_later
BibTeX
@misc{fu2026vynn,
author = {Fu, Zanwen},
title = {{VYNN} {AI}, Six Months Later: A Simpler Agent and Stricter Numbers},
year = {2026},
month = oct,
howpublished = {\url{https://zanwenfu.com/blog/vynnai_six_months_later}},
note = {Published October 7, 2026}
}Published October 7, 2026.