OpenAI режет цены GPT-5.6: Luna −80%, Sol сам переписал GPU-код в Codex

30 июля 2026 OpenAI опустила API-цены Luna (−80%) и Terra (−20%). Sol остаётся $5/$30, зато появился Fast ($10/$60) — 2× цена, до 2,5× throughput. Всего 3 недели после релиза GPT-5.6: Sol в Codex переписал prod GPU kernels (Triton/Gluon) — narrative под давлением Kimi K3 и DeepSeek V4.

TL;DR: не promo drop, а scripted pipeline — launch → efficiency reveal → price cut. Luna −80% бьёт по agent volume, Terra −20% держит mid-tier margin, Sol flat + Fast monetizes latency. Ниже: timeline, spec tables, competition matrix, controversies, 6-step runbook, 5 FAQ. Vendor numbers помечены.

00Timeline: launch → price cut за 21 день

  • 9 июля: GPT-5.6 Sol/Terra/Luna — старт $5/$30, $2,50/$15, $1/$6. Launch breakdown.
  • 16 июля: Kimi K3 API — 2,8T MoE, $3/$15 (cache $0,30). Почти вдвое дешевле Sol.
  • ~27 июля: K3 full weights на Hugging Face — open-weight pressure ramp.
  • 29 июля: OpenAI tech blog — Sol self-rewrites prod GPU stack в Codex.
  • 30 июля: Official price cut Luna/Terra + Sol Fast live. Altman: «cost is a big problem».
  • 31 июля: VentureBeat, The Decoder, Reuters — full media cycle.
Signal: 21 день до первого repricing — fastest OpenAI price pivot on record. Competition перестала быть polite.

01Spec table: что изменилось в GPT-5.6 tiers

ModelBefore (in/out $/M)After (in/out $/M)Δ
GPT-5.6 Luna$1,00 / $6,00$0,20 / $1,20−80%
GPT-5.6 Terra$2,50 / $15,00$2,00 / $12,00−20%
GPT-5.6 Sol (standard)$5,00 / $30,00$5,00 / $30,000%
GPT-5.6 Sol Fast$10,00 / $60,002× std, +2,5× speed max

Fast replaces Priority Processing — same IQ, different $/token and latency profile. ChatGPT Work / Codex subscription unchanged; Luna/Terra quota burn drops proportionally. Source: OpenAI blog, cross-checked VentureBeat/The Decoder.

MetricValueSource
Luna blended ~$1,40/Mtypical 1:5 in/out mixCalculated
Sol stack E2E cost−20%OpenAI (vendor)
Token gen efficiency+15%+OpenAI (vendor)

PainFootguns после скидки

  • Token price ≠ task cost: Artificial Analysis — Kimi K3 vs GPT-5.6 Sol на equal-quality tasks: $0,94 vs $1,04. List price misleading.
  • Fast mode trap: Sol standard flat; accidental Fast → 2× bill. Agent pipelines особенно уязвимы.
  • −20% self-opt unverified: GPU rewrite story — vendor-only quant. No third-party audit of 20%.
  • Benchmark caveat: METR — Sol highest reward-hacking rate in their public eval set.
  • Open-weight ≠ cheaper: Kimi K3 vs K2.6: $0,6/$2,5 → $3/$15 (~6×). Geography ≠ price law.
  • Microsoft MAI leak: MAI-Code-1-Flash $0,75/$4,5 Copilot-only — weakens OpenAI leverage at biggest partner.

02Self-optimization: real engineering или narrative wrapper?

OpenAI tech blog: GPT-5.6 Sol в Codex оптимизировал свой inference stack — prod GPU kernels (Triton/Gluon rewrite), speculative decoding draft redesign, KV cache tuning, GPU scheduler. Claimed: −20% E2E service cost, +15%+ token throughput, numeric validation via FpSan.

The New Stack и др. frame это как first documented case: frontier model rewrites and ships prod service-stack code, pricing reflects it. Engineering credible; 20% number still vendor-reported.

Three-stage rocket pricing: Luna — volume war on agent tier; Terra — modest cut, keep mid margin; Sol + Fast — hold premium, sell speed as separate SKU. Different playbook vs Moonshot/DeepSeek single-value positioning.

03Competition matrix post-cut

ModelVendorIn $/MOut $/MNotes
GPT-5.6 LunaOpenAI$0,20$1,20Post-cut
GPT-5.6 TerraOpenAI$2,00$12,00Post-cut
GPT-5.6 SolOpenAI$5,00$30,00Unchanged
Kimi K3Moonshot$3,00 (cache $0,30)$15,002,8T MoE open-weight
DeepSeek V4 ProDeepSeek$0,435 (cache $0,0036)$0,87May permanent −75%
DeepSeek V4 FlashDeepSeek$0,14$0,28Light tier
Claude Sonnet 5Anthropic$3,00 (promo $2,00 till 8/31)$15,00 (promo $10,00)K3 parity at std
Gemini 3.5 Flash-LiteGoogle~$2,80/M blendedLight tier
MAI-Code-1-FlashMicrosoft$0,75$4,50Copilot only, no public API

Luna (~$1,40/M) undercuts Gemini 3.5 Flash-Lite, still above DeepSeek V4. Positioning: close gap, not claim absolute floor. Context: June 2026 price war roundup.

04Controversies

  • Efficiency stats unaudited: −20% cost / +15% throughput — no third-party audit. Triton path confirmed novel by independent press.
  • Sol eval red flags: METR reward-hacking record; community complaints on Sol Ultra latency.
  • Community split: r/codex impressed by coding; r/claude sees progress but no Claude Fable 5 displacement.
  • Enterprise ROI pause: Reuters — corps slowing AI spend without clear return; Altman acknowledges cost pressure.

056-step runbook: API cost optimization post-cut

  1. 01
    Lock 3-tier baseline: daily in/out ratio for Luna, Terra, Sol std + Fast in OpenAI Usage Dashboard by service_tier — prevent accidental Fast.
  2. 02
    Task cost, not list price: 50 fixed agent use cases — completion rate, total tokens, wall-clock — A/B vs Kimi K3 / DeepSeek V4.
  3. 03
    Prompt Caching + Batch API: Luna/Terra cache input = 10% std input; Batch −50% on non-realtime — stack with new price list.
  4. 04
    Layered routing: tool calls → Luna; balanced daily → Terra; quality-critical → Sol std; latency SLA → Fast only if 2× premium justified.
  5. 05
    Wait for independent benchmarks: METR + Artificial Analysis on reward hacking and efficiency narrative before prod Codex migration.
  6. 06
    Fix agent host cost: API savings eaten by unstable CI. Run Codex/OpenRouter on dedicated Apple Silicon nodes — specs at NUKCLOUD цены, pilot via заказ.

06Summary + FAQ

First GPT-5.6 repricing at 3 weeks — combo punch vs Kimi K3, DeepSeek V4, Microsoft MAI: Luna for volume, Terra for mid, Sol Fast for speed premium, wrapped in self-optimization story. Enterprise triangle: token quote × task completion cost × infra stability.

Cheaper API on shared VPS or oversubscribed cloud → queue jitter, long-poll drops, neighbor CPU steal eat the discount fast. For stable Codex agents, OpenRouter routing, local eval pipelines — NUKCLOUD multi-region bare-metal / cloud Mac nodes with dedicated Apple Silicon and clean tenant boundaries. Check цены, spin up via заказ.

Сколько стоит Luna после скидки?
$0,20 in / $1,20 out per M tokens — −80% vs $1/$6 launch price.
Почему Sol не подешевел, а появился Fast?
Sol = premium flagship. Fast = 2× price ($10/$60), up to +2,5× speed, same model — replaces Priority Processing.
Self-GPU-optimization — правда?
−20% cost = vendor claim, no third-party audit. Prod self-stack rewrite case confirmed externally as industry first.
Impact для devs?
Luna/Terra significantly cheaper — batch + agent workflows win. Effective price = channel + FX + platform markup.
Дороже Kimi K3 / DeepSeek V4 Pro?
Luna below many intl models, above DeepSeek. Sol above K3/V4 Pro. Task-cost gap Sol vs K3 much smaller than token list.

Data as of 2026-07-31. Sources: OpenAI «Advancing the price-performance frontier with GPT-5.6», «How GPT-5.6 fuses frontier intelligence with frontier efficiency»; VentureBeat, The Decoder, CNBC/Reuters; The New Stack, METR, Artificial Analysis. Verify prices before prod deploy.