Each dot is a model at release, plotted by the blended cost at your selected workload mix. Two years of releases, one chart. The 2026 cluster contains GLM-5.1 sitting on top of Claude Opus 4.7 in Elo space at a fraction of the price.
Holding capability constant, blended price has fallen by roughly two orders of magnitude. Adjust the workload mix above to see how the gap changes for your actual use case.
Math is computed live from your selected workload mix. The doubling time and rate update accordingly, so you can see the deflation curve as it actually applies to your spend.
Notice the ratio column. Anthropic and OpenAI charge 5× to 6× more for output than input. Chinese open models tend toward 3× to 4×. This means the cost gap is widest for output-heavy workloads and narrowest for RAG-heavy. Toggle the workload mix above to see it move.
The pricing-strategy tell
Anthropic prices output at 5× input. Chinese open models sit closer to 3-4×. Output token generation is more expensive in compute terms, but not by that much. The 5× ratio captures a brand premium, not just compute. The Chinese pricing is closer to actual serving cost.
Workload mix matters more than headline
A pure RAG workload sees the open-vs-closed economics narrow because input pricing converges across providers. An agentic or code-gen workload sees the gap widen because that is where the frontier premium is hidden.
Routing implication
For input-heavy workblocks, the open-vs-closed economics still favor open but the urgency is lower. For output-heavy generation and agent loops, open-weight routing is closer to a 25-30× cost saving than 21×. Build the routing layer with workload-aware pricing.
Watch over next 6 months
GLM-5.1's Opus parity claim verifying on independent SWE-Bench and stable LMArena vote counts.
Whether Anthropic compresses the input:output ratio. A move from 5× to 4× would signal pricing pressure from below.
GLM-5.1 listing on Bedrock and Azure Foundry. Each addition tightens the procurement noose.
DeepSeek V5 timing. Mid-2026 release with another step-function is base case.
Whether the deflation rate at your workload mix holds or starts decelerating toward 3-5×.