OpenWeightsTerminal

GLM 5.3 Flash

GLM 5.3 Flash

RTX PRO 6000 + 2× GB10 · EXL3 4bpw · TensorFold

Decode · tok/s
120
Prefill
—
tok/100W
20
GLM-5.3-Flash (EXL3 4bpw) on TensorFold, fully local: 80 → 120 tok/s (+50%) by adding one GPU over plain 2.5GbE. Layers 0-26 on an RTX PRO 6000; layers 27-44 + head on 2× GB10.

First-person hybrid box. Rank is the quoted 120 with thinking off. 80 is the 2× GB10 baseline in the same sentence, not a range.

Officialclaimed
Open pair