Astra. Fable. Grok. The names sound expensive. The engineering question is whether the extra intelligence earns its bill.
I think of the labs as players in a repeated game, with two strategies worth watching. Crown chasers push for the highest capability they can demonstrate. Their flagship has to do something the competition cannot. Price breakers attack the cost of capable models. Their threat is a model good enough to make an expensive incumbent uncomfortable.
These are competitive strategies. One lab can play both. A flagship can chase the crown while its little brother attacks the bill.
When the little brother makes a task cheaper, the flagship has to defend its premium. A lab might cut the price, push capability further (like loosening guardrails to widen access to capabilities), or win on something the cheaper model cannot do. Or win on more reliable tool use or lower latency. Every move changes what the next player has to beat.
My bet is that this pressure is a powerful driver of progress. After reviewing the data on the Frontier class and the Pareto class, I can observe capability at the Pareto class that did not exist in March of this year. Pareto models are getting better per dollar.
Frontier models and Pareto models
Crown chasers and price breakers describe how the labs compete. Frontier and Pareto describe what we buy. Here is the vocabulary I use.
Frontier model. A model at the leading edge of demonstrated general-purpose capability, at a given date. It earns consideration for the difficult work that weaker models cannot yet do reliably.
Pareto model. The efficient model choice for a specified workload. Among the configurations evaluated, it meets the quality, reliability, and response-time requirements at the lowest total cost.
The frontier is a moving capability boundary. The Frontier Model Forum uses the term broadly for the current state of the art. I look for evidence across difficult reasoning, engineering, and tool use. A model can keep every capability it has today and lose its frontier position when a rival advances. The unlock with a frontier model is the degree to which it can challenge or improve a technical decision, or contribute an insight for the greater whole. Call it higher-order thinking: reasoning through the consequences of a decision and navigating the branches of possibility.
Our Pareto definition adds a selection rule to Pareto efficiency. Pareto efficiency identifies tradeoffs: improving one objective means giving up another. Our workload supplies the quality bar and the evaluation meter. Evaluations establish which configurations clear the bar. Among those, we choose the lowest total cost, and retries and human repair count in that bill.
I run such selections by starting at the frontier and testing downwards. The strongest model establishes what a good result looks like. Each cheaper model then tests how well I have specified the work and equipped the agent.
A failure at a cheaper model can mean two things. The model is too weak, or the workload is underspecified: vague instructions, missing context, a brittle environment. A frontier model tends to reason its way around those defects, at a cost, and hides them. A cheaper model exposes them. Fixing what it exposes improves the whole system, so testing downwards is a quality step as well as a cost step.
These labels can overlap. A frontier model may also be Pareto-efficient, and may be our cheapest reliable choice, all depending on the nature of the task. Model family, size, and open or closed weights do not settle that question. Neither does a low token price. The comparison belongs to a configuration doing a job. Anthropic's model-selection guide makes the same practical case for testing capability, speed, cost, and reasoning effort against actual work. Anthropic presents two recommended paths: efficiency-first, or capability-first. I believe that starting capability-first, then progressing downwards, is the more generally applicable choice. It lets you find the specific definition of the system more closely than starting bottom up.
The good news is that the number of models delivering near-frontier intelligence at Pareto costs is growing. Comparing early September with March shows how quickly the truly competent Pareto-class candidates have multiplied.
March: the challenger box is empty
Take the Artificial Analysis leaderboard captured on 6 March 2026. Among 138 configurations with a measured score and usable cost, Gemini 3.1 Pro Preview held the highest score.
The challenger box sets two conditions for taking on the leader:
Score: at least 80% of the leader’s score.
Cost: at most 20% of the leader’s bill for running the same index.
Meet both, and a configuration is an 80/20 challenger. Near the leading score. A fraction of the bill.
In March, nobody qualified.
1 Gemini 3.1 Pro Preview. Leading score. 57.2 index, $892.28 index cost.
2 Gemini 3 Flash Preview (Reasoning). 81.2% of the leading score, 31.19% of the leader’s cost. Cheapest efficient configuration at 80% of the leader.
3 MiMo-V2-Flash (Feb 2026). 72.5% of the leading score, 7.65% of the leader’s cost. Cheapest efficient configuration at 70% of the leader.
The staircase is the Pareto frontier: configurations for which no other option scores at least as well for less, or better for the same cost. March had efficient choices. None cleared our 80/20 threshold. A place on the staircase and a place in the challenger box are separate qualifications.
September: 29 challengers
In the saved 4 September snapshot, Claude Fable 5.1 at max effort held the highest score. Apply the same relative rule and 29 of 129 eligible configurations sit inside the challenger box.
Different reasoning efforts count as separate configurations. Each of these 29 challengers clears the same score floor and cost ceiling. The leader and the evaluation mix changed between the two snapshots, so this compares each market against its own leader. It does not measure the price decline for a fixed level of capability.

1 Claude Fable 5.1 (max). Leading score. 65.7 index, $8,523.16 index cost.
2 GLM-5.3-Flash. 87.5% of the leading score, 1.62% of the leader’s cost. Cheapest efficient configuration at 80% of the leader.
3 GPT-5.6 Luna (high). 71.5% of the leading score, 0.66% of the leader’s cost. Cheapest efficient configuration at 70% of the leader.
Three points on the staircase show the whole argument. Claude Fable 5.1 at max effort is the frontier. GLM-5.3-Flash is the challenger: 87.5% of the leading score for 1.6% of the leader’s index cost. GPT-5.6 Luna at high effort sits at the bottom of the staircase, the cheapest configuration that keeps at least 70% of the leading score, at 0.66% of the cost. All three are Pareto-efficient. Which one is the Pareto model depends on where the workload sets its bar.
GLM-5.3-Flash makes the challenger concrete, and the team behind it has been building toward this for years. Z.ai, formerly Zhipu AI, grew out of Tsinghua University. Nathan Lambert’s history of GLM at Interconnects traces the company’s founding in 2019, the first GLM research in 2021, and ChatGLM in 2023. His reading of the current generation is that it is a post-training story, not a distillation story: the same base model, with more environments and more compute spent training on them. This is a sustained research program showing up in the price breakers’ game.
GLM-5.3-Flash has 320 billion parameters, with 18 billion active per token, and combines sparse and linear attention. It takes text, images, and video as input. Those are specific architectural choices behind a model with an unusually small benchmark bill. They earn it a place in the challenger box. Whether it becomes my Pareto model depends on the work it can deliver.
The task sets the budget
Fable is the reliable engineer in my setup. When designing a new system, working through an unfamiliar codebase, or deciding what to build, I want to maximize the intelligence I can get. A better model might find the simpler architecture, catch a broken assumption, or save me a day on the wrong problem. That is specifically what the frontier premium gets me. The opportunity cost of building a worse system far outweighs the cost of the maximum intelligence that might prevent it.
Opportunity cost is the value of the best alternative we give up. In model selection, that might be a better design, an hour of engineering time, or budget we could put to work elsewhere.
On the contrary, if the model is part of a system, the calculus changes. Inside an engineered system, the model works with constraints and specific instructions. The model receives a contract to follow, and novelty is a defect. Inside such a system, evaluations are run on the model at scale to reach a probabilistic conclusion that the model is suitably instructed, with errors prevented through retry loops, for example. Once a cheaper model meets that contract reliably, the costly flagship has little to no justified presence.
My default is Frontier for discovery, Pareto for iterated delivery. Pay for judgement while the problem is open. Once the problem is defined and a cheaper model is proven, settle for the cheapest option that meets the bar.
What comes after the price breakers?
Over the next year, I expect frontier models to take on longer design and engineering jobs from an incomplete brief, and to become more able to correct their own course and explore decision branches. For Pareto models, I expect greater speed and a lower price per task.
These are roles in a system. The same lab, and sometimes the same model at different reasoning settings, may serve both.
Working forecast, September 2026. I expect frontier models to improve judgement and decision making, and Pareto models to expand the work we can afford to repeat.
Recent pricing changes show why neither class will have one simple price:
Anthropic estimates that Fable 5.1 will cost 25% less than Fable 5 on typical token-billed workloads because of cheaper cache reads, with savings up to approximately 45% for highly agentic work. Fable 5.1 announcement.
Claude Fable 5.1 through Anthropic's Message Batches API costs $5 per million input tokens and $25 per million output tokens. Exactly 50% of its standard $10 and $50 rates. Requests run asynchronously and can take up to 24 hours. Fable 5.1 pricing, Message Batches API.
DeepSeek announced peak and off-peak billing effective 16 August 2026, with off-peak rates half the peak rates. DeepSeek pricing update.
OpenAI lists Sol’s promotional pricing as available at least through 21 November 2026. Google Cloud lists introductory standard global Gemini 3.7 Flash rates of $0.75 input and $3.75 output per million tokens through December, with $1.50 and $7.50 scheduled for 1 January 2027. OpenAI pricing, Google Cloud pricing.
GPT-6 Astra prices context length and urgency separately. Standard rates are $10 per million input tokens and $50 per million output tokens. A request with more than 272,000 input tokens is billed at $20 and $75 for the whole request. Batch and Flex cost half the applicable standard rates. Fast mode costs twice them. Astra pricing.
I expect model selection to become more tightly tied to the type of task the model is applied to. Reusable context, a flexible delivery window, or the end of a promotion can change the cost of a task. A stronger frontier model can earn a larger bill if it saves equivalent human work.
Model selection can thus be consolidated into creation versus iteration. A Frontier model is the creator, the Pareto model the iterator.
Sources and measurement
The figures use the 6 March leaderboard capture and the Artificial Analysis leaderboard retrieved on 4 September 2026. These are fixed observations, not a live feed.
The chart dataset, March observations, September observations, and pricing source notes are available as JSON. The chart dataset records source URLs, file hashes, and inclusion counts.
The definition and batch-pricing notes record the additional sources checked through Exa on 6 September 2026. The publication source notes cover the Fable 5 redeployment, the GLM history and architecture, and the Astra pricing. The workload definitions are our editorial choices. The batch rates are Anthropic's published prices.
Both figures are also available as images, with the three highlighted points labelled: March as PNG or SVG, September as PNG or SVG. The highlights follow one rule on both dates: the leader, then the cheapest efficient configuration keeping 80% of its score, then the cheapest keeping 70%.
Each point is a non-deprecated configuration with a non-estimated score and a positive recorded cost for running that snapshot's full index. We exclude configurations without those measurements. The reference is the highest-scoring eligible configuration, with lower cost breaking ties. Both plots use the same relative axes.
The score percentage is a fraction of a benchmark index, not a fraction of intelligence. Cost includes the evaluation's token usage at the recorded prices. It is not a quote for a production workload. See the Artificial Analysis methodology.
Two September costs were calculated from Artificial Analysis token counts and external prices: Muse Spark 1.3 (max) from Meta, and Nex-N2-Pro from OpenRouter. Their saved records identify that calculation. Motif 3 and K2 Horizon 375B remain excluded because the saved work did not establish a public price.



