DeepSeek V4.1 Flash Enters Beta Testing, Reaching Speeds Up to 507 Tokens/s

nashnova research
今天发布阅读约 5 分钟

On September 8 DeepSeek opened beta access to V4.1 Flash; developers clocked output speeds up to 507 tokens/s on a new architecture with native multimodal support, priced the same as V4 Flash, available for just two days.

01

How fast is it, really?

Developers on X report a peak output speed of 507 tokens/s.
Several testers ran the same benchmark — an SVG pelican-riding-a-bicycle animation — and logged 300–328 tokens/s consistently.
This means → even ignoring the outlier, the stable range is an order of magnitude above most mainstream models. Speed is no longer the bottleneck; consistency is.
02

What's new, and how do you try it?

V4.1 Flash uses a new model architecture with native multimodal support — text, images, and other inputs handled by a single model.
Developers keep their existing base_url and swap the model name to `deepseek-v4.1-flash-expires-on-0910`. Each account is capped at 20 concurrent requests.
In plain terms = no API migration needed — change one field and go. But the model name carries its own expiry date: auto-offline on September 10, a two-day window.
03

Has the price changed?

Beta pricing matches V4 Flash exactly: off-peak input (cache hit) ¥0.05 per million tokens, input (cache miss) ¥1.5, output ¥4.5.
Peak-hour prices double: input at ¥0.1 and ¥3, output at ¥9.
This means → DeepSeek is not charging a speed premium — more capability, same price. But this is beta pricing; whether it holds at general release is unconfirmed.
04

What does this release cadence signal?

DeepSeek has shipped updates in rapid succession: V4 Flash on July 31, V4 Pro on August 13, V4 Flash Vision Exp on August 21 (open-sourced August 31).
V4.1 Flash beta follows less than two weeks later.
This reflects a high-frequency release cycle — model generations are turning over faster. Whether the beta's speed numbers carry over to the production release remains to be seen.

市场有风险,内容仅供研究参考,不构成投资建议。