There is no best model. The frontier is jagged. Advancements are constant, surprising, and uneven.
Model advancements are uneven, surprising, and accelerating - just last week, Gemini 3.8 Flash, Muse Spark 1.3, Fable 5.1, and GPT-6 Astra were all released. Every model has unique characteristics (or personalities, if you will) and they excel at different things. As such, no single model will be the best model for everything. But a system that takes advantage of each model’s unique strengths and mitigates its weaknesses can outperform any one model.
The frontier is jagged
The headline leaderboards models are measured against reduce their capabilities to a single dimension and, over time, make the frontier look like a staircase that’s slowly - or rather, very quickly - being climbed.
However, vertical AI companies understand that their domains don’t reduce to singular dimension. Rather they understand that models are multifaceted: They have different characteristics, personalities, and they excel at different tasks. Some models need very specific instructions, while others can handle ambiguity. Some are quick to ask for help/clarification, while some are reluctant to ask at all. Some models are great at long-context reasoning, some are good at manipulating documents, and some are good at instruction-following for writing styles.
And this dimension is only one of several. Models differ in cost and latency by a factor of fifty or more. Certain providers are very reliable while others often fail either generally or in odd, specific ways. On top of that there are differences in geographic availability and governance. This makes the frontier of intelligence jagged: different models drive the frontier in different aspects.
The frontier moves quickly and unevenly
Not only is the frontier jagged, it’s constantly changing. New providers are popping up with new models and even new versions of existing models often have vastly different characteristics than previous versions. And it looks like the rate of change is increasing. One year ago there were for all intents and purposes two labs: OpenAI and Anthropic. Now SpaceXAI, Meta Superintelligence, and Google look like serious competitors. Not to mention that Chinese labs like Z.ai and Moonshot have caught up rapidly with frontier closed-weight labs.
The palette of potential models has never been more diverse or faster changing. Just last week, Gemini 3.8 Flash, Muse Spark 1.3, Fable 5.1, and GPT-6 Astra were all released. Frontier leaderboards are changing weekly, and the pareto frontier is being pushed outwards along its entire curve: tiny, efficient models are being released for specialized workflows or certain intelligence-saturated tasks, while big models push the ceiling of intelligence and reasoning. And inference efficiency gains and/or open weights competition leads to frequent steep price cuts: recently OpenAI slashed prices by 80% for their Luna model.
The best model is all the models
Because models are different and the model landscape is changing so quickly, the best way to ensure optimal outputs for our users is to make sure the best possible model does any given job. Often “best” means that it completes the task at a high level of quality at the lowest possible cost and latency. Achieving this generally means decomposing the problem into constituent parts that different models tackle with their own context windows and specialized tools.
Each model has strengths and weaknesses. A multi-model system can trace the outer envelope of them all - taking the best capability in each direction to create a whole that outperforms any individual model.
Vertical AI companies are uniquely positioned to map the nuanced model map and spot “model arbitrage” opportunities - places where changing models gives more performance, lower cost, lower latency or the gain of other desired properties - that aren’t immediately clear from a public leaderboard. They can do this because they have two very valuable things: in-house expertise and valuable real-world usage.[1]
Expertise creates high-quality evals at the level of “updating change-of-control clause in poorly-formatted Italian shareholder’s agreement” not just “good at legal”. Having in-house expertise means vertical AI companies can continuously revise evals and add more difficult, detailed cases as models improve and climb the scores.

Real-world usage ensures that what’s being created by experts matches what’s valued by actual users. Without live usage, evals are best-guess, and value is bound to be left on the table. Live usage gives insights into long-tail failures and use cases that experts and product teams won’t consider a priori and creates a good environment to validate hypothesis and roll out improvements.
The two feed each other. Usage shows where evals lack coverage; in-house experts turn gaps into detailed evals and failures into regressions as developers make harness tweaks to squeeze all possible performance out of each model. With any new model release, it runs against the suite, takes over the tasks where it wins, and a controlled rollout reveals new nuances to sharpen our understanding and advance the frontier.
Bet on the system
Every model on the board will be superseded, probably sooner rather than later. The harness will be changed as capabilities expand. The systematic analysis and model selection persists between it all - the expertise that encodes wanted behavior, the live usage that sharpens it.
There is no best model.





