OpenWeightsTerminal

GLM 5.3

GLM 5.3

4x DGX Spark · Int4/Int8 · sparkDash

Decode · tok/s
30
Prefill
840
tok/100W
12.5
Full GLM-5.3 (753B) TP4 on 4× DGX Spark. Prose decode: 30.0 tok/s at c1 (sparkDash), 24.9 tok/s with thinking on (RigMark). Prefill: 840 at 32k. Int4/Int8 mixed weights, native MTP (K=2).

First-person 4× Spark recipe. Rank is the quoted 30.0 thinking-off prose. Tilde prefill and the 4-stream aggregate stay off the rank.

Officialclaimed
Open pair