All notes
InterviewZH → EN

Fireworks co-founder Benny Chen on open models, token growth, and the eval gap

Benny Chen (Fireworks AI) · 硅谷坐标 Silicon Valley Vector

Token charts are the vanity metric of open-source AI; revenue is the honest one. A 45-minute tour of where open models actually stand, from the platform that says it serves more open-model tokens per day than Gemini and OpenAI publish for their B2B APIs — and an argument that the scarcest skill in the industry right now is writing evals, not training models.

Open vs closed models

  • Open models are closing the gap faster than even the believers expected. His team came out of PyTorch at Meta convinced that open source eventually catches up in every valuable software layer — it did for operating systems, it did for databases. What nobody priced in was this year’s pace, at a fraction of closed-model prices, while closed-model revenue kept compounding anyway.
  • Distillation is a bridge, not a moat. It can bootstrap a model to a strong level quickly, but the API leakage that made it easy is being patched, and several strong open models barely relied on it. The durable engine is clear, well-funded evals plus falling data costs.
  • The closed labs’ real long-term asset is buyer psychology, not capability. Like iPhone versus Android: the gap narrows to trivia, yet price-insensitive customers keep paying for the leader, and enterprise procurement stays conservative — nobody gets fired for buying the incumbent. He gives that advantage two or three years, and declines to call five.

Vertical agents vs the AGI track

  • Building a top vertical agent — Harvey in legal, Doximity and Heidi in medical — and chasing AGI are different jobs. One is fine-grained evals and data cleaning for a niche; the other is proving you can solve Riemann-class problems. Jagged intelligence is a feature for a vertical SaaS company and a bug for an AGI lab, which is why a frontier lab crushing any single vertical would not pay for its valuation. The tell: last year’s frontier releases washed out many customer fine-tunes, this year almost none.
  • Almost everything is being reframed as a coding problem. Once even PowerPoint and Excel are expressed as code, the entire existing stack of coding tools and RL recipes applies, and a vertical application can be tuned quickly. What does not become coding tends to become deep research; most workloads sort into one of the two.
  • The industry over-serves the STEM benchmarks and under-serves everyone else. Slides, documents, content research, video — the humanities workload — may be the bigger market, since most work ultimately serves people. It is also why computer-use agents stalled after the initial awe: plain text is cheaper, so workflows migrated to text, and the category revives as VLM prices fall.

Enterprise adoption and customization

  • The scale check: Fireworks discloses roughly 40–50 trillion tokens a day of open-model serving — more, he says, than the public B2B API figures from Gemini and OpenAI combined suggest for closed equivalents.
  • But watch revenue, not tokens. Fireworks grew from a $100M to a $1B run rate while Anthropic grew far faster; free traffic pollutes the router leaderboards, and a frontier token counts the same as a budget one. Buyers pay for the absolute best, so a customized model’s job is beating the frontier on the task — not winning on price-performance.
  • POCs die of missing evals. Even in 2026, teams whose AI spend is a major cost centre still vibe-test. His pitch: writing evals is to AI products what writing unit tests was to SaaS — the same muscle, transitioned — and one rigorous eval suite can settle an open-vs-closed procurement by itself. People who can do this are scarce.
  • Who should customize: vertical SaaS companies, whose customer base lets them build evals that cover the real distribution — not individual end-organizations, who are better off encoding their workflow into skills and internal tools. Customized RL needs around a thousand environments, not millions of examples, because rollouts multiply the data; once evals are pinned down, rebasing onto each new open-source release is fast.
  • Deployment is tilting toward the cloud. The security-and-trust objections are solvable, and multi-tenancy is decisive: few enterprises can saturate their own GPUs, while aggregated demand can — so as AI becomes a real line item, sharing the cost beats owning the hardware.

Inference, incentives, and infrastructure

  • Most Fireworks traffic is customized models, not vanilla base models. Serving many customers’ divergent workloads surfaces failure patterns a single-model first-party API never sees, and that breadth feeds back into finer-grained optimization — the moat is the distribution of workloads, not any one trick.
  • Incentive design as strategy: consulting and pure-training shops get paid once, so their incentives end at delivery. Fireworks earns on customers’ ongoing inference traffic — it only wins if the customized model keeps delivering in production. The endgame competitor is not other neoclouds but the CSPs, who can run the whole chain; hence partnering tightly instead (first-party on Azure).
  • The binding constraint for the next year or two is compute at a reasonable price. Revenue is compute-bound, and he expects the industry to stay supply-constrained for at least another year.
  • If open models catching up crashes chip stocks, the logic is backwards: stronger open models weaken pricing power at the model layer, which favors the infrastructure underneath it — the same confusion as the market’s reaction when DeepSeek first landed.
Watch on YouTube