Inside the Cut: Why Luna Fell 80% but Terra Only 20%
OpenAI's July 30 announcement cuts GPT-5.6 Luna's API price by 80%, from $1.00 to $0.20 per million input tokens and $6.00 to $1.20 per million output tokens [2]. Terra, the mid-tier model, gets a comparatively modest 20% reduction, from $2.50 to $2.00 per million input tokens and $15.00 to $12.00 per million output tokens [2]. Sol, the flagship, keeps its $5/$30 per-million-token pricing untouched, but gains a new Fast Mode that trades money for speed: up to 2.5x faster responses for twice the price, with zero change to the model's underlying intelligence [3].
OpenAI ties the broader move to internal engineering gains: GPT-5.6 reportedly helped optimize the company's own production and GPU-serving code, cutting end-to-end serving costs by roughly 20% and improving token-generation efficiency by more than 15% through techniques like speculative decoding and kernel optimization [1]. The company has not detailed why Luna's cut is four times steeper than Terra's on a percentage basis, leaving the tier-specific math open to interpretation. What's certain is that the same lower per-token pricing now also counts against ChatGPT Work and Codex subscription quotas, meaning paying subscribers effectively get more usage for the same monthly fee [2].



