OpenWeightsTerminal

GLM 5.3 Flash

GLM 5.3 Flash

4× RTX PRO 6000 · EXL3 · TensorFold 0.6.2

Decode · tok/s
264
Prefill
—
tok/100W
26.4
4x RTX PRO 6000 (250 W per GPU): 264 tok/s on prose, single stream

Aevonix single-host recipe on Mia EXL3. Rank is the quoted 264 prose one-stream.

Officialclaimed
Open pair