OpenWeightsTerminal
Pairs

Qwen3.8-Flash-Next

2× DGX Spark · FP8 · vLLM

Decode · tok/s
58
Prefill
—
tok/100W
24.2
code at 65 tok/s on 1 spark vs 58 tok/s on my 2 … my vllm serve runs the official fp8 across both sparks with mtp on

Same post as the 65 TensorFold pair. Rank is the quoted 58 code on two Sparks. 44.5 is the 128k decode in the follow-up. 20 to 30 on GLM Flash is a range, not a pair.

Officialclaimed
Open pair