Decode · tok/s
58
Prefill
—
tok/100W
24.2
code at 65 tok/s on 1 spark vs 58 tok/s on my 2 … my vllm serve runs the official fp8 across both sparks with mtp on
Same post as the 65 TensorFold pair. Rank is the quoted 58 code on two Sparks. 44.5 is the 128k decode in the follow-up. 20 to 30 on GLM Flash is a range, not a pair.
