Umans DeepSeek V4 Flash (lab) Playground Retired
playground experiment; the text V4 Flash lab window ended when the vision lab opened
retired Aug 11, 2026 · weights ↗
370.9tok/s
throughput · p50 · whole period
1.84s
TTFT · p50 · whole period
100.00%
uptime · whole period
DeepSeek V4 Flash as a Labs experiment, open for a short test window: temporary, not a permanent id. DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release, on a 1M-token context. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Access is seat-gated through the Labs page while an experiment is live. It is offered at limited capacity and availability, so expect it to be flaky and to go down under load: crash it, give it a moment, and try again. When the window ends, the model keeps serving as the pay-per-token umans-deepseek-v4-flash-0731.
Trends
Speed over its final 90 days
peak 421.1 tok/s · Aug 4final 380.3 tok/s
Aug 4, 2026retired Aug 11, 2026Aug 13, 2026
final 908ms
Aug 4, 2026retired Aug 11, 2026Aug 13, 2026
Changelog
Events for Umans DeepSeek V4 Flash (lab)
Aug 32026
The V4 Flash lab continues as umans-deepseek-v4-flash-0731-lab Testing
The V4 Flash pilot closed at the pay-per-token release: the production id now bills per token, so the seat-gated pilot on it ended rather than charge anyone by surprise. The lab reopened on the new umans-deepseek-v4-flash-0731-lab id with a smaller cohort - free, seat-gated, same experimental capacity as before. The model keeps serving as umans-deepseek-v4-flash-0731 regardless: the model stays. (The text lab closed on 2026-08-11 when the vision lab opened - see the next entry.)