Online users
The number of people the deployment serves at the same time.
What your team gets
See the service capacity at a glance, then choose the profile that fits your workload.
The number of people the deployment serves at the same time.
How quickly each user's prompt, files, and history enter the model.
How quickly each user sees the answer arrive.
The combined Prefill and Decode capacity for the whole team.
The amount of code, documents, and history one user can keep in a session.
Native FP4 inference
From a compact two-node team to a 64-user production service
A single DGX Spark is not a delivery profile for V4 Flash 0731. The service catalog starts at two 128 GB nodes, then scales to four nodes with IB and to a full PRO 6000 node.
| Deployment profile | Online users | Prefill / user | Decode / user | System Prefill | System Decode | Context / user | Cost index |
|---|---|---|---|---|---|---|---|
| 2× DGX SparkMinimum 1M-context delivery | 4–8 @ 1M | 217.38–434.75 | 16.13–32.27 | 1,739 | 129.06 | 1M | 0.018× |
| 4× DGX Spark + IBRecommended small-team profile | 8–12 @ 1M | 194.0–291.0 | 17.40–26.09 | 2,328 | 208.74 | 1M | 0.040× |
| 2× RTX PRO 60008-user · approx. 400K context profile | 8 | 1,606.13 | 72.59 | 12,849 | 580.7 | ≈400K | 0.45× |
| 4× RTX PRO 6000Balanced production profile | 8 @ 1M | 1,669.0 | 100.56 | 13,352 | 804.5 | 1M | 0.45× |
| 4× RTX PRO 6000High-concurrency profile | 64 @ 1M | 208.63 | 39.70 | 13,352 | 2,540.5 | 1M | 0.45× |
| 8× RTX PRO 6000Full-node · 1M long-context profile | 64 @ 1M | 421.9 | 80.34 | 27,001.6 | 5,141.9 | 1M | 0.45× |
| 4× RTX PRO 6000DIE measured · 1M profile | 8 @ 1M | 1,380.3 | 83.16 | 11,042.4 | 665.3 | 1M | 0.30× |
| 8× RTX PRO 6000DIE measured · 1M long-context profile | 64 @ 1M | 355.7 | 67.73 | 22,764.8 | 4,334.7 | 1M | 0.30× |
单位均为 tok/s;每用户速度按系统总吞吐除以所列在线用户数计算。上下文表示为每位用户预留的容量,不代表每次吞吐测试都将提示词填满至该长度。2× PRO 6000 的 8 用户档是唯一低于 1M 的配置,约为每用户 400K;其余档位均按每用户完整 1M 容量规划。
Hardware cost and delivery capacity
One eight-GPU H200 SXM server is the 1.00× hardware cost baseline. Every card names its model and active-user profile: H200 uses V4 Pro 0813 FP8; Blackwell and Spark use V4 Flash 0731.
Performance for every user
DeepSeek V4 Flash 0731, DeepSeek V4 Pro 0813, and Kimi K3 are shown as complete service profiles with per-user speed, total throughput, online capacity, context, and system cost.
同时在线用户 · 1M 长上下文生产档
同时在线用户 · 1M 长上下文高性价比档
同时在线用户 · 1M 上下文 · 640GB KV 池
同时在线用户 · 1M 上下文 · 512GB KV 池
并发请求 · 800K 容量 · 8K/1K 负载测速
同时在线用户 · 每用户完整 1M 上下文
同时在线用户 · 每用户完整 1M 上下文
并发智能体任务 · 1M 上下文容量
1M 表示每位用户可使用的上下文容量,不代表每次测速请求都使用了完整 1M-token 提示词。KV 显存开销因模型而异,V4 Flash、V4 Pro 与 Kimi K3 分别规划。
Context capacity
Long context keeps code, documents, tool results, and working history together. The service profile tells you how many complete sessions fit on the node.
The eight-GPU PRO 6000 node has 768 GB of physical VRAM; the 6000D node has 672 GB. Resident weights, runtime workspace, and the model-specific KV pool are budgeted independently so both delivery profiles reserve 1M context capacity for 64 users. The compact catalog starts at 2× DGX Spark with 4–8 users at 1M capacity.
After weights and runtime allocation, the profile keeps about 640 GB as pooled context memory: 80 users × 8 GB each at 1M. The 6000D profile serves 64 users at the same 1M context.
The 32-GPU profile serves two complete 1M users; the 48-GPU profile raises full-context capacity to three.
Choose your starting point
Reference baseline
8× H200 SXM · vLLM · MTP-2 · 256 concurrent requests · 8K input / 1K output · system Prefill 10,147.3 / Decode 1,268.7 tok/s. Hardware cost index: 1.00×.
以上数据来自特定测试环境,仅供采购参考,具体以正式合同与 POC 实测结果为准。
DeepAgent IE
Tell us the model, number of users, and context requirement. We will return a deployment profile with a clear performance and hardware cost index.