OpenAI announced aggressive price cuts for GPT-5.6 models: Luna dropped 80% to $0.20 per million input tokens (from ~$1.20), Terra fell 20% to $2/$12 input/output pricing, and GPT-5.6 Sol introduced a new 2.5× faster mode at 2× standard price with no intelligence degradation. The cuts are tied to systems-level efficiency improvements across the model, inference stack, and agentic harness layer, including autonomous kernel optimization and speculative decoding gains.
The most striking metric: GPT-5.4 full (OpenAI's March flagship) at its peak benchmark score (51 on Arena Elo equivalent) now costs 13× more than GPT-5.6 Luna at the same performance level. Over just four months, OpenAI has compressed the cost curve by an annualized rate of ~2000× when holding intelligence constant. Downstream, ChatGPT's auto-review and Codex CLI are migrating from GPT-5.4 to Luna, yielding roughly 10× cost savings per task.
For AI infrastructure buyers and LLM platform teams: this pricing shock signals a fundamental decomposition in what 'model cost' means. Distillation, speculative decoding, inference optimization, and prompt caching are driving the marginal unit economics below what system-level efficiency alone would predict. As open models (Poolside's Laguna, Thinky's Inkling, DeepSeek) improve, the proprietar advantage shifts from raw capability to orchestration—harness design, tool routing, and context compaction that OpenAI can tune end-to-end.