Anthropic's Claude Sonnet 5 is priced at introductory $2/$10 per million input/output tokens through August 31, 2026, after which standard pricing rises to $3/$15—a 50% price increase on both axes. The pricing cliff means teams budgeting on the assumption of continued Sonnet 5 sub-$2/$10 rates have one month to lock in contracts or migrate workloads. Opus 5, meanwhile, remains fixed at $5/$25, narrowing the price-to-performance gap and making Opus increasingly competitive for complex agentic work relative to the mid-tier Sonnet.
In parallel, OpenAI announced an 80% price cut on GPT-5.6 Luna, reducing it from $1.00 to $0.20 per million input tokens. The cut funded not by margin compression but by engineering: Sol (GPT-5.6's full model) autonomously rewrote its own production GPU kernels and speculative-decoding pipeline inside Codex, the first confirmed instance of a frontier model self-optimizing its serving stack. OpenAI framed this as the beginning of a feedback loop: as models improve and operate more autonomously, the company's ability to compress serving costs itself accelerates.
For practitioners: the dynamics are diverging. Anthropic is raising Sonnet 5 pricing but holding Opus steady; OpenAI is collapsing inference costs via algorithmic efficiency. The winner depends on your workload: if you run high-volume classification or routing on Sonnet, budget the 50% hike. If you're heavy on Luna for low-stakes inference, the price drop makes it the cheapest frontier-model option in production. Both moves signal that the era of stable LLM pricing is over—model providers are using algorithmic innovation and pricing elasticity as competing levers to capture token volume and margin simultaneously.