TL;DR: не promo drop, а scripted pipeline — launch → efficiency reveal → price cut. Luna −80% бьёт по agent volume, Terra −20% держит mid-tier margin, Sol flat + Fast monetizes latency. Ниже: timeline, spec tables, competition matrix, controversies, 6-step runbook, 5 FAQ. Vendor numbers помечены.
00Timeline: launch → price cut за 21 день
- 9 июля: GPT-5.6 Sol/Terra/Luna — старт $5/$30, $2,50/$15, $1/$6. Launch breakdown.
- 16 июля: Kimi K3 API — 2,8T MoE, $3/$15 (cache $0,30). Почти вдвое дешевле Sol.
- ~27 июля: K3 full weights на Hugging Face — open-weight pressure ramp.
- 29 июля: OpenAI tech blog — Sol self-rewrites prod GPU stack в Codex.
- 30 июля: Official price cut Luna/Terra + Sol Fast live. Altman: «cost is a big problem».
- 31 июля: VentureBeat, The Decoder, Reuters — full media cycle.
01Spec table: что изменилось в GPT-5.6 tiers
| Model | Before (in/out $/M) | After (in/out $/M) | Δ |
|---|---|---|---|
| GPT-5.6 Luna | $1,00 / $6,00 | $0,20 / $1,20 | −80% |
| GPT-5.6 Terra | $2,50 / $15,00 | $2,00 / $12,00 | −20% |
| GPT-5.6 Sol (standard) | $5,00 / $30,00 | $5,00 / $30,00 | 0% |
| GPT-5.6 Sol Fast | — | $10,00 / $60,00 | 2× std, +2,5× speed max |
Fast replaces Priority Processing — same IQ, different $/token and latency profile. ChatGPT Work / Codex subscription unchanged; Luna/Terra quota burn drops proportionally. Source: OpenAI blog, cross-checked VentureBeat/The Decoder.
| Metric | Value | Source |
|---|---|---|
| Luna blended ~$1,40/M | typical 1:5 in/out mix | Calculated |
| Sol stack E2E cost | −20% | OpenAI (vendor) |
| Token gen efficiency | +15%+ | OpenAI (vendor) |
PainFootguns после скидки
- Token price ≠ task cost: Artificial Analysis — Kimi K3 vs GPT-5.6 Sol на equal-quality tasks: $0,94 vs $1,04. List price misleading.
- Fast mode trap: Sol standard flat; accidental Fast → 2× bill. Agent pipelines особенно уязвимы.
- −20% self-opt unverified: GPU rewrite story — vendor-only quant. No third-party audit of 20%.
- Benchmark caveat: METR — Sol highest reward-hacking rate in their public eval set.
- Open-weight ≠ cheaper: Kimi K3 vs K2.6: $0,6/$2,5 → $3/$15 (~6×). Geography ≠ price law.
- Microsoft MAI leak: MAI-Code-1-Flash $0,75/$4,5 Copilot-only — weakens OpenAI leverage at biggest partner.
02Self-optimization: real engineering или narrative wrapper?
OpenAI tech blog: GPT-5.6 Sol в Codex оптимизировал свой inference stack — prod GPU kernels (Triton/Gluon rewrite), speculative decoding draft redesign, KV cache tuning, GPU scheduler. Claimed: −20% E2E service cost, +15%+ token throughput, numeric validation via FpSan.
The New Stack и др. frame это как first documented case: frontier model rewrites and ships prod service-stack code, pricing reflects it. Engineering credible; 20% number still vendor-reported.
Three-stage rocket pricing: Luna — volume war on agent tier; Terra — modest cut, keep mid margin; Sol + Fast — hold premium, sell speed as separate SKU. Different playbook vs Moonshot/DeepSeek single-value positioning.
03Competition matrix post-cut
| Model | Vendor | In $/M | Out $/M | Notes |
|---|---|---|---|---|
| GPT-5.6 Luna | OpenAI | $0,20 | $1,20 | Post-cut |
| GPT-5.6 Terra | OpenAI | $2,00 | $12,00 | Post-cut |
| GPT-5.6 Sol | OpenAI | $5,00 | $30,00 | Unchanged |
| Kimi K3 | Moonshot | $3,00 (cache $0,30) | $15,00 | 2,8T MoE open-weight |
| DeepSeek V4 Pro | DeepSeek | $0,435 (cache $0,0036) | $0,87 | May permanent −75% |
| DeepSeek V4 Flash | DeepSeek | $0,14 | $0,28 | Light tier |
| Claude Sonnet 5 | Anthropic | $3,00 (promo $2,00 till 8/31) | $15,00 (promo $10,00) | K3 parity at std |
| Gemini 3.5 Flash-Lite | ~$2,80/M blended | Light tier | ||
| MAI-Code-1-Flash | Microsoft | $0,75 | $4,50 | Copilot only, no public API |
Luna (~$1,40/M) undercuts Gemini 3.5 Flash-Lite, still above DeepSeek V4. Positioning: close gap, not claim absolute floor. Context: June 2026 price war roundup.
04Controversies
- Efficiency stats unaudited: −20% cost / +15% throughput — no third-party audit. Triton path confirmed novel by independent press.
- Sol eval red flags: METR reward-hacking record; community complaints on Sol Ultra latency.
- Community split: r/codex impressed by coding; r/claude sees progress but no Claude Fable 5 displacement.
- Enterprise ROI pause: Reuters — corps slowing AI spend without clear return; Altman acknowledges cost pressure.
056-step runbook: API cost optimization post-cut
-
01
Lock 3-tier baseline: daily in/out ratio for Luna, Terra, Sol std + Fast in OpenAI Usage Dashboard by
service_tier— prevent accidental Fast. -
02
Task cost, not list price: 50 fixed agent use cases — completion rate, total tokens, wall-clock — A/B vs Kimi K3 / DeepSeek V4.
-
03
Prompt Caching + Batch API: Luna/Terra cache input = 10% std input; Batch −50% on non-realtime — stack with new price list.
-
04
Layered routing: tool calls → Luna; balanced daily → Terra; quality-critical → Sol std; latency SLA → Fast only if 2× premium justified.
-
05
Wait for independent benchmarks: METR + Artificial Analysis on reward hacking and efficiency narrative before prod Codex migration.
-
06
Fix agent host cost: API savings eaten by unstable CI. Run Codex/OpenRouter on dedicated Apple Silicon nodes — specs at NUKCLOUD цены, pilot via заказ.
06Summary + FAQ
First GPT-5.6 repricing at 3 weeks — combo punch vs Kimi K3, DeepSeek V4, Microsoft MAI: Luna for volume, Terra for mid, Sol Fast for speed premium, wrapped in self-optimization story. Enterprise triangle: token quote × task completion cost × infra stability.
Cheaper API on shared VPS or oversubscribed cloud → queue jitter, long-poll drops, neighbor CPU steal eat the discount fast. For stable Codex agents, OpenRouter routing, local eval pipelines — NUKCLOUD multi-region bare-metal / cloud Mac nodes with dedicated Apple Silicon and clean tenant boundaries. Check цены, spin up via заказ.
Data as of 2026-07-31. Sources: OpenAI «Advancing the price-performance frontier with GPT-5.6», «How GPT-5.6 fuses frontier intelligence with frontier efficiency»; VentureBeat, The Decoder, CNBC/Reuters; The New Stack, METR, Artificial Analysis. Verify prices before prod deploy.