旗舰私有推理引擎

DeepAgent IE —— 企业私有原生 FP4 推理引擎

把 DeepSeek V4 Flash 0731、DeepSeek V4 Pro 0813 与 Kimi K3 部署在企业自有基础设施中。按团队真正能感知的数字选择:同时在线、每用户速度、系统总吞吐、上下文容量和硬件成本。

DeepAgent IE icon
实测性能DeepAgent IE
5,141.9system Decode tok/s · Flash · 64 users @ 1M
12,849cold Prefill tok/s · 2× PRO 6000
348system Decode tok/s · V4 Pro · 80 users
1Mcontext per user · V4 Pro / Kimi K3

What your team gets

Five numbers make the decision simple

See the service capacity at a glance, then choose the profile that fits your workload.

01

Online users

The number of people the deployment serves at the same time.

02

Prefill per user

How quickly each user's prompt, files, and history enter the model.

03

Decode per user

How quickly each user sees the answer arrive.

04

System throughput

The combined Prefill and Decode capacity for the whole team.

05

Context per user

The amount of code, documents, and history one user can keep in a session.

Native FP4 inference

DeepSeek V4 Flash 0731

A single DGX Spark is not a delivery profile for V4 Flash 0731. The service catalog starts at two 128 GB nodes, then scales to four nodes with IB and to a full PRO 6000 node.

Deployment profileOnline usersPrefill / userDecode / userSystem PrefillSystem DecodeContext / userCost index
2× DGX SparkMinimum 1M-context delivery4–8 @ 1M217.38–434.7516.13–32.271,739129.061M0.018×
4× DGX Spark + IBRecommended small-team profile8–12 @ 1M194.0–291.017.40–26.092,328208.741M0.040×
2× RTX PRO 60008-user · approx. 400K context profile81,606.1372.5912,849580.7≈400K0.45×
4× RTX PRO 6000Balanced production profile8 @ 1M1,669.0100.5613,352804.51M0.45×
4× RTX PRO 6000High-concurrency profile64 @ 1M208.6339.7013,3522,540.51M0.45×
8× RTX PRO 6000Full-node · 1M long-context profile64 @ 1M421.980.3427,001.65,141.91M0.45×
4× RTX PRO 6000DIE measured · 1M profile8 @ 1M1,380.383.1611,042.4665.31M0.30×
8× RTX PRO 6000DIE measured · 1M long-context profile64 @ 1M355.767.7322,764.84,334.71M0.30×

单位均为 tok/s;每用户速度按系统总吞吐除以所列在线用户数计算。上下文表示为每位用户预留的容量,不代表每次吞吐测试都将提示词填满至该长度。2× PRO 6000 的 8 用户档是唯一低于 1M 的配置,约为每用户 400K;其余档位均按每用户完整 1M 容量规划。

Hardware cost and delivery capacity

Put every deployment on one cost scale

One eight-GPU H200 SXM server is the 1.00× hardware cost baseline. Every card names its model and active-user profile: H200 uses V4 Pro 0813 FP8; Blackwell and Spark use V4 Flash 0731.

V4 Pro 0813 · 256 users

H200 SXM

1.00×
System
1× 8-GPU server
Prefill / Decode
10,147.3 / 1,268.7 tok/s
Hardware cost index
1.00×
V4 Flash 0731 · 64 用户

RTX PRO 6000

0.45×
System
1× 8-GPU server
Prefill / Decode
27,001.6 / 5,141.9 tok/s
Hardware cost index
0.45×
V4 Flash 0731 · 64 用户

RTX PRO 6000D

0.30×
System
1× 8-GPU server
Prefill / Decode
22,764.8 / 4,334.7 tok/s
Hardware cost index
0.30×
V4 Flash 0731 · 4–8 用户

DGX Spark

0.018×
System
1× two-node cluster
Prefill / Decode
1,739 / 129.06 tok/s
Hardware cost index
0.018×

Performance for every user

Give every user enough performance

DeepSeek V4 Flash 0731, DeepSeek V4 Pro 0813, and Kimi K3 are shown as complete service profiles with per-user speed, total throughput, online capacity, context, and system cost.

Native FP4 · IE measured

DeepSeek V4 Flash 0731

64

同时在线用户 · 1M 长上下文高性价比档

Hardware
8× RTX PRO 6000D
Total VRAM
672 GB
Online users
64 @ 1M
Prefill / user
355.7
Decode / user
67.73
System Prefill
22,764.8
System Decode
4,334.7
Hardware cost index
0.30×
Native FP4 · IE measured

DeepSeek V4 Pro 0813

64

同时在线用户 · 1M 上下文 · 512GB KV 池

Hardware
16× RTX PRO 6000D
Total VRAM
1,344 GB
Online users
64 @ 1M
Prefill / user
29.45
Decode / user
3.75
System Prefill
1,884.8
System Decode
240.0
Hardware cost index
0.60×
H200 reference · FP8

DeepSeek V4 Pro 0813

256

并发请求 · 800K 容量 · 8K/1K 负载测速

Hardware
8× H200 SXM
Total VRAM
1,128 GB
Online users
256
Prefill / user
39.64
Decode / user
4.96
System Prefill
10,147.3
System Decode
1,268.7
Hardware cost index
1.00×
Native FP4 · IE measured

Kimi K3

2 / 3

同时在线用户 · 每用户完整 1M 上下文

Hardware
32 / 48× RTX PRO 6000D
Total VRAM
2,688 / 4,032 GB
Online users
2 / 3 @ 1M
Prefill / user
682.92 / 731.84
Decode / user
100.77 / 107.12
System Prefill
1,365.8 / 2,195.5
System Decode
201.5 / 321.4
Hardware cost index
1.20× / 1.80×
H200 reference · FP4

Kimi K3

3

并发智能体任务 · 1M 上下文容量

Hardware
32× H200 SXM
Total VRAM
4,512 GB
Online users
3
Prefill / user
797.2424
Decode / user
6.0352
System Prefill
2,391.73
System Decode
18.11
Hardware cost index
4.00×

1M 表示每位用户可使用的上下文容量,不代表每次测速请求都使用了完整 1M-token 提示词。KV 显存开销因模型而异,V4 Flash、V4 Pro 与 Kimi K3 分别规划。

Context capacity

Keep the whole job in one private session

Long context keeps code, documents, tool results, and working history together. The service profile tells you how many complete sessions fit on the node.

DeepSeek V4 Flash 073164

1M users · 8× PRO 6000 / 6000D

The eight-GPU PRO 6000 node has 768 GB of physical VRAM; the 6000D node has 672 GB. Resident weights, runtime workspace, and the model-specific KV pool are budgeted independently so both delivery profiles reserve 1M context capacity for 64 users. The compact catalog starts at 2× DGX Spark with 4–8 users at 1M capacity.

DeepSeek V4 Pro 081380

1M users · 16× PRO 6000

After weights and runtime allocation, the profile keeps about 640 GB as pooled context memory: 80 users × 8 GB each at 1M. The 6000D profile serves 64 users at the same 1M context.

Kimi K32 / 3

1M users · 32 / 48 GPUs

The 32-GPU profile serves two complete 1M users; the 48-GPU profile raises full-context capacity to three.

Choose your starting point

Start with the workload, then scale with confidence

01

Small private team

2× DGX Spark
Recommended model
V4 Flash 0731
Online profile
4–8 users @ 1M
System Decode
129.06 tok/s
Cost index
0.018×
03

高性价比生产

8× RTX PRO 6000D
Recommended model
V4 Flash 0731
Online profile
64 users @ 1M
System Decode
4,334.7 tok/s
Cost index
0.30×
04

旗舰长上下文

16 / 32× RTX PRO 6000
Recommended model
V4 Pro / Kimi K3
Context / user
1M / 1M
Online profile
80 / 2 users
Cost index
0.90× / 1.80×

Reference baseline

H200 SXM · DeepSeek V4 Pro FP8

8× H200 SXM · vLLM · MTP-2 · 256 concurrent requests · 8K input / 1K output · system Prefill 10,147.3 / Decode 1,268.7 tok/s. Hardware cost index: 1.00×.

H200 SXM 参考数据

以上数据来自特定测试环境,仅供采购参考,具体以正式合同与 POC 实测结果为准。

DeepAgent IE

Turn private hardware into usable AI capacity

Tell us the model, number of users, and context requirement. We will return a deployment profile with a clear performance and hardware cost index.