OpenWeightsTerminal

GLM 5.3 Flash

GLM 5.3 Flash

2x DGX Spark · EXL3 · TensorFold

Decode · tok/s
85
Prefill
—
tok/100W
35.4
GLM-5.3-Flash on 2x DGX Spark with TensorFold: 85 tok/s single stream, 174 tok/s total across 8 streams (code). Screenshot: GLM-5.3-Flash-EXL3, x1 stream 85.0 tok/s.

First-person 2× Spark. Rank is the quoted 85.0 single-stream code. 174 / 173.6 is the 8-stream aggregate, not a pair. bpw is not in the quote.

Officialclaimed
Open pair