Notes on the transactions we make, the sectors we watch, and the conviction behind both.
Envoy's view: task volume will fragment far faster than economic value. Frontier models will lose routine work while retaining disproportionate importance wherever being moderately better creates a discontinuous jump in value.
The state of the market: a Ferrari closely tailed by a Mustang
Unlike proprietary models such as Claude and ChatGPT, open-weight models (often also called “open source”) publicly release the parameters that determine how they operate. Historically, they faced two limitations. First, they resembled a free engine rather than a finished car: the core technology cost nothing to download, but users still had to specialize, integrate, host it, and maintain it. Second, even after that effort, the engine materially underperformed the factory-built alternatives from Anthropicand OpenAI.
That is changing quickly. Two consequential open-weight models shipped in mid-July, most notably Kimi K3 from China's Moonshot. K3 scored within four points of the leading closed model on Artificial Analysis's Intelligence Index. Soon after, Jensen Huang argued that open models are necessary to make AI “economically sustainable.” Today, two of the top six models on cost-performance are open-weight.
The free engine is now powerful enough to justify building the chassis around it.
The challenge to frontier models is clear: Open-weight models are becoming good enough to absorb an increasing proportion of workloads, commoditizing the layer of intelligence just below it on a continuous basis. Think of a next-best alternative that stays right on your tail. While the frontier has taken the overwhelming share of economic rent from the development of LLMs to date, if the layer immediately beneath it is increasingly capableandclose to free, the future value extraction potential of the frontier itself comes into question.
We have spent the last several months unpacking what follows, working through public data, sell-side research, and conversations with the people building, buying and deploying these systems. Below we lay out our findings, how this may play out in the medium term, and the implications for us as investors.
The capability gap is real, but so is the price gap
Open-weight models remain behind the absolute frontier. Artificial Analysis scores Claude Opus Max, the leading closed model, at 61, against 57 for Kimi K3. A year ago the leading open-weight model scored 22 against 35 for the leading closed model, a 13-point gap. Two things follow. Regarding the gap, open weights are closing the distance rapidly, and yet the remaining four points still matter. In practical terms the outputs feel like the difference between the smartest kid in the class and the fourth smartest; both clear the bar consistently.
But as the task set becomes longer, harder and less forgiving, the difference becomes more apparent, and more decisive at the tail.
Regarding the rate of improvement, today's best open-weight models are roughly twice as capable as last year's frontier models. Whatever last year's models already did well (document summarization, contract drafting, data cleansing), today's open models do better.
How and why is this happening? Capabilities are cheaper to copy than to invent. Frontier labs bear the enormous cost of discovering what works; open developers increasingly reproduce much of that performance through distillation, synthetic data and improved training techniques. The frontier pays the discovery cost; open models increasingly pay only the replication cost.
The relative economics thus also favor open weights. On Artificial Analysis's weighted cost per completed task, Claude Opus Max costs $2.03 against $0.72 for Kimi K3.

Widen the frame and the asymmetry gets starker: DeepSeek V4 Pro scores 44 at four cents a task, one fiftieth the price of the leader for roughly two thirds of its measured capability.
This raises the obvious CFO question: where am I paying 3 to 50 times more than I need to, for a difference in performance that does not matter? In what instances am I sending a Ferrari to the grocery store when a Prius would do?

The volume shift is early, and it will be significant
Most organizations we have spoken with, including sophisticated Silicon Valley technology companies, remain early in shifting workloads to open-weight models.
One Fortune 500 CTO captured the state of the market well. Twelve months ago his organization rewarded “tokenmaxxing,” the idea that whoever spends the most on models deserves credit for embracing AI. Six months ago that changed. The company had burned through its annual AI budget in the first three months of the year. His teams now run a homegrown router that compares task-specific accuracy, speed and price.
Today that router sends roughly 30% of tasks to open-weight models, a proportion he expects to grow. However, he also believes, as do several others we spoke with, that the strongest frontier models remain meaningfully better at complex coding and agentic work: planning, reading unstated intent, and figuring out what to do next.
Published work points the same direction. A RouteNLP study earlier this year found routers send roughly one quarter to one third of requests to frontier models. A trained router in the Berkeley LMSYS work delivered approximately 95% of frontier benchmark quality while routing only a quarter of calls to the frontier. Intelligent routing makes selective dependence on the frontier possible.

We do not expect a single open-weight model to capture enormous volume on its own. We expect hundreds or thousands of sufficiently capable models, with sophisticated routing choosing among them at the task level. Cursor’s agentic coding product moved meaningful workloads from Anthropic to smaller specialized models, illustrating how responsive routing can drive fragmentation without eliminating reliance on frontier intelligence.
OpenRouter sits directly at that point of selection. Stripe's reported interest in acquiring it for as much as $10 billion underscores the value of choosing which model handles a task, at what price and under what constraints.
Why frontier models still capture disproportionate value
Anthropic and OpenAI reportedly run above $60 billion and $40 billion of annualized revenue respectively as of this writing. Moonshot, whose K3 sits four points off the top of the intelligence index, reportedly reached $300 million of ARR in June. The combined difference runs roughly three hundred to one. Why?
While the revenue delta partially reflects business model differences, the extreme dispersion has more to do with how customers value frontier intelligence versus every layer below it. Our hypothesis has two parts. First, reliability at the edge. On Artificial Analysis's knowledge-and-hallucination measure, the leading open-weight models score between -10 and +6. The leading closed model scores +20. A four-point gap on the composite conceals a far larger gap in whether the model makes up facts, follows an entire chain of reasoning, catches a flawed intermediate assumption, or recovers from an unexpected result.
Second, value scales discontinuously at the extreme.
Consider three-point shooting in the NBA. A 39% shooter and a 33% shooter are separated by six makes per hundred attempts, a difference that looks almost trivial measured literally. Economically it can be enormous. The 39% shooter bends the defense, creates space for everyone else, and likely commands $50 million a year. A player whose principal value is shooting at 33% plays little and fights for a $4 million contract. Six percentage points do not produce six percent more value. They determine whether the skill matters at all. At the extremes, tiny differences are all that matter.
Frontier intelligence has the same structure. A four-point benchmark advantage does not make every answer four percent better. It means one model clears a threshold the other does not:it finds the exception that changes the conclusion, reconciles the conflicting evidence instead of averaging past it, or carries a complex analysis far enough to produce a defensible answer rather than a plausible-looking one. Itcan be the difference between an output you can use and one that ships a fatal flaw.
Structural dynamics also amplify seemingly small leads. The best lawyer or engineer has finite hours, which leaves ample economic room for the second, tenth and hundredth best practitioner. A model carries no such constraint. If one model is even modestly preferred, every buyer can choose it at once. Preference compounds from there: buyers do not re-benchmark every release, so the leading model accumulates usage, tooling, evaluations and institutional habit. A small technical lead becomes a large commercial one.
We believe our hypothesis survives under several future conditions. For example, in the near term, frontier labs may again widen their lead—due tocapital-intensive synthetic data development and post-training, which creates a larger pool of high value tasks that only they can handle reliably. This does not change thatopenweights hold a structural advantage in overall volume share due to continuous improvement, and wherever buyers care mostabout unit cost, customization and control over deployment.
What this means for how we invest
Frontier models will be fine. Open-weight models will take share. The more interesting question is where value accrues between those two truths.
We believe it accrues to the companies that make fragmented intelligence useful: inference providers that deliver structurally better economics, such as Baseten (deepest into control plane) and Sail Research (long horizon agent durability); post-training platforms that make specialization cheaper and more accessible such as Databricks Mosaic and Thinking Machines' Tinker; and routing layers that continuously choose the right model for each task either standalone or a core differentiator within AI native applications.
These may not remain distinct categories. Inference providers sit naturally positioned to absorb routing and post-training into a broader control plane, while standalone routers may struggle to defend a thin aggregation layer. The companies worth owning will do more than pass requests through. Fal.ai is a company we think fits well. Fal is simultaneously the OpenRouter of media (benefits from fragmentation) and owns proprietary inference margin on workloads where optimization is a real, defensible edge. We heard from several customers that its inference engine runs diffusion models ~4x faster than generic clouds with clear ROI.
The same opportunity exists at the application layer. The best applications will become domain-specific routers themselves, combining frontier and open intelligence with proprietary data, evaluations and workflows to deliver an outcome no underlying model provides on its own. ElevenLabs increasingly looks like this: it owns critical audio models and the customer workflow while orchestrating multiple language models and providers underneath. Its agent platform already offers multi-model access and automated failover.
That sharpens the test we apply when we underwrite. If a competitor can post-train an open model next year and arrive at the same product, the durable value is small. The companies that endure will own the workflow, generate proprietary feedback from usage, and continuously arbitrage a changing model landscape on behalf of their customers.
The model layer will fragment. The value will accrue to those best positioned to orchestrate it.