← all models
Umans DeepSeek V4 Pro Retired
umans-deepseek-v4-pro-0813 · DeepSeek
no longer available; use umans-deepseek-v4.1-flash
retired Sep 14, 2026 · replaced by umans-deepseek-v4.1-flash · weights ↗
Retired
115.5tok/s
throughput · p50 · whole period
1.28s
TTFT · p50 · whole period
100.00%
uptime · whole period

DeepSeek V4 Pro, served from the official 0813 release: DeepSeek's flagship coding and reasoning MoE, built for long-horizon agentic coding and demanding tool-heavy workloads on a 1M-token context window. The 0813 release supersedes the April preview with substantially stronger agentic performance. Reasoning has three modes: non-think (none), think high (high, the default) and think max (max). Billed per token ($1.32 / $3.96 / $0.044 per 1M; input / output / cache read). Served on our own GPU infrastructure with high availability. Deprecated: use `umans-deepseek-v4.1-flash` instead (sunset 2026-09-14).

Jun 18, 2026retired Sep 14, 2026Sep 16, 2026
Trends

Speed over its final 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
Jun 18, 2026retired Sep 14, 2026Sep 16, 2026
TTFT p50 · time to first token, lower is better
Jun 18, 2026retired Sep 14, 2026Sep 16, 2026
Changelog

Events for Umans DeepSeek V4 Pro

incl. gateway-wide announcements
Sep 142026
Retired: Umans DeepSeek V4 Pro Retired
umans-deepseek-v4-pro-0813 reached its sunset and was retired in favour of DeepSeek V4.1 Flash: it leaves the catalog and the model pickers, and its history stays on the past-models list. Requests still pinned to the old id keep routing for a short grace tail - migrate to umans-deepseek-v4.1-flash (the DeepSeek lineage's successor: 1M context, native vision) now; a final cutoff date will be announced ahead of the hard stop.
Older events 1
Aug 152026
Released pay-per-token: Umans DeepSeek V4 Pro Released
umans-deepseek-v4-pro-0813 joins the lineup as the long-context coding flagship: DeepSeek's official 0813 release of V4 Pro, the checkpoint that served here as the seat-gated pre-release lab since August 13, with a 1M context window and thinking at high effort by default (dial to max when a task deserves more). Billed per token: $1.32 / $3.96 / $0.044 per 1M (input / output / cache read). It succeeds GLM 5.2, which is deprecated and sunsets on August 23, 2026. The pre-release window's metrics stay on the model's status page as its pre-release period (before Aug 15). Served on our own GPU infrastructure with high availability.