OpenWeightsTerminal

DeepSeek v4.1 Flash

DeepSeek v4.1 Flash

RTX PRO 6000 · Q2 · llama.cpp b11412

Decode · tok/s
57.7
Prefill
—
tok/100W
9.6
使用GPUは、RTX PRO 6000。DeepSeek-V4-Flash-Q2-0731 … 生成 57.70 t/s … GPURAM 84.72 GB。llama.cpp(build 11412)、500文字までの日本語生成。

First-person llama.cpp bench. Rank is the quoted 57.70 generation t/s. Filename says V4-Flash Q2-0731.

Officialclaimed
Open pair