OpenWeightsTerminal

GLM 5.3 Flash

GLM 5.3 Flash

2× DGX Spark · EXL3 · TensorFold

Decode · tok/s
60.4
Prefill
1950
tok/100W
25.2
Serve the full 1M-token context window with 4 concurrent requests on dual DGX Sparks

The author's sparkDash run through the OpenAI API, not reproduced by the desk. Checkpoint Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw (official weights). Rank uses the single-stream prose figure. The DFlash2 drafter is CC BY-NC-ND 4.0, non-commercial.

Officialclaimed
Open pair