GPT-6 Inference Speed Up 50%, OpenAI Has Already Shifted 80-90% of Research to GPT-7 and GPT-8
nashnova research
OpenAI boosted GPT-6's default inference from 30 to 50 tokens per second, yet 80–90% of its research now targets GPT-7 and GPT-8 — the speed bump is a short-term patch; the next generation is the real bet.
50% faster — can users actually feel it?
GPT-6 Astra and Sol now generate 50 TPS, up from 30, across all ChatGPT logged-in products and third-party tools like OpenCode and Devin.
The underlying tokenizer was swapped at the same time, compressing total tokens needed for the same task. This means → it is not just "faster typing" — the same job now costs fewer tokens to finish.
Some developers report the real-world gain feels limited; hitting a stable 50 TPS requires the API channel. In plain terms = the headline number looks good, but actual experience varies by use case.
80–90% of resources on the next generation — so what is GPT-6?
Applied-research head Boris Power disclosed that 80–90% of research resources are now focused on GPT-7, GPT-8, and beyond.
Continued GPT-6 optimization is internally viewed by some as an "extremely short-sighted" near-term spend. This means → OpenAI believes many current pain points will dissolve once the next-gen models ship — today's patches have limited shelf life.
This reflects a deeper stance: OpenAI would rather let the current product be "good enough" and save real engineering firepower for the generational leap.
"Build strong, then distill" — how do big and small models relate?
OpenAI's product playbook: train the most capable large model first, then use it as a "teacher" to transfer abilities — via distillation — into smaller, cheaper models.
Case in point: distilling GPT-4o into GPT-4o mini lifted accuracy on a specific task from 64.67% to 79.33%, nearly matching GPT-4o's own 79.67%.
In plain terms = the budget model you use daily has a top-tier model coaching it behind the scenes — small models are cheap, but their capability comes from big-model research.
Same $200 a month — how much is Claude worth vs. ChatGPT?
SemiAnalysis tested both by subscribing, running heavy coding and agent workloads to the weekly cap, then converting usage to equivalent API pricing. At the $200/month tier, ChatGPT Pro delivered roughly $2,084 in equivalent API value; Claude Max 20x delivered roughly $11,726 — a gap of more than 5×.
The ratio held across tiers: 5.6× at $200, 5.4× at $100, 5.6× at $20.
This means → for heavy coding and agent workloads, Claude's premium subscription delivers far more usable capacity dollar-for-dollar than ChatGPT's equivalent plan.
The P&L behind subscriptions — who is losing money?
Anthropic breaks even on Claude Pro and Max 5x only when users hit roughly 20% utilization; OpenAI's ChatGPT Plus and Pro 5x start losing money once utilization exceeds about 11.4%.
Last week OpenAI cut the effective quota on its $200 ChatGPT Pro top-tier plan by roughly half; SemiAnalysis links this directly to the margin pressure above.
Put simply = the more power users OpenAI attracts, the deeper the losses — the quota cut is not stinginess; it is triage. Whether GPT-6's speed gains can narrow the value gap with Claude, and whether GPT-7 and GPT-8 can deliver the generational leap Power describes, will be the core tests of OpenAI's competitiveness in the next phase.
市场有风险,内容仅供研究参考,不构成投资建议。
