OpenWeightsTerminal

Qwen3.8-27B

Qwen3.8-27B

RTX PRO 6000 Blackwell Max-Q · NVFP4 · TensorFold

Decode · tok/s
272.8
Prefill
—
tok/100W
109.1
同じQwen3.8-27B NVFP4重みのdecodeはTensorFold 272.8/vLLM 151.4 tok/s。RTX PRO 6000 Blackwell Max-Q・250W、1件・code greedy・256出力。下書き方式は異なります。

Quoted code-greedy decode on one PRO 6000 Max-Q. Draft method differs from the vLLM cell in the same sentence, so the two are not a locked comparison.

Officialclaimed
Open pair