OpenAI Cuts Prices on GPT-5 and Two Other Models

Miles Bennett
Published todayAbout 6 min read

OpenAI slashed prices on two GPT-5.6 models just three weeks after launch, with the lightest version, Luna, dropping roughly 80% — a pace driven by relentless competition from low-cost Chinese open-source models.

01

How much did each model drop?

GPT-5.6 Luna, the fastest and cheapest tier, fell about 80% — now $0.20 per million input tokens and $1.20 per million output tokens.
GPT-5.6 Terra, the mid-tier model, came down 20% to $2 input / $12 output per million tokens.
The flagship GPT-5.6 Sol was not included. This means → OpenAI is cutting at the light and mid tiers while holding premium pricing at the top.
02

Why cut prices just three weeks in?

OpenAI's official reason: improved serving efficiency — the same intelligence now costs less to run.
The more pressing force is competition from low-cost Chinese open-source models, which keep pulling down the market's price expectations for AI calls.
In plain terms = OpenAI is not being generous — competitors have lowered the ceiling, and standing still means losing developers.
03

What does an 80% Luna cut actually change?

Luna targets lightweight call scenarios — simple Q&A, text processing, high-volume but low-complexity tasks.
An 80% price drop sharply reduces the marginal cost of these calls, making the tier far more attractive to small and mid-size developers.
This means → OpenAI's playbook is clear: use Luna's ultra-low price to capture the developer entry point, pull users into the ecosystem, then funnel high-value demand toward Terra and Sol.
04

What is shifting in the industry's pricing logic?

Both OpenAI and Anthropic have adjusted pricing, rate limits, and usage policies multiple times — cuts are becoming routine, not exceptional.
The driver: next-generation reasoning models consume massive token volumes during long-running agentic tasks, making users extremely price-sensitive per token.
This reflects a structural shift — AI models are moving from "call-by-call" usage to "continuous runtime," and pricing must follow, from high-unit-price / low-volume toward low-unit-price / high-volume.

Content is for reference only, not financial advice.

OpenAI Cuts Prices on GPT-5 and Two Other Models · nashnova