One quoted capture. Rank only lives inside a filter.
GLM 5.3 Flash
RTX PRO 6000 + 2× GB10 · EXL3 4bpw · TensorFold
Decode · tok/s
120
Prefill
—
tok/100W
20
GLM-5.3-Flash (EXL3 4bpw) on TensorFold, fully local: 80 → 120 tok/s (+50%) by adding one GPU over plain 2.5GbE. Layers 0-26 on an RTX PRO 6000; layers 27-44 + head on 2× GB10.
First-person hybrid box. Rank is the quoted 120 with thinking off. 80 is the 2× GB10 baseline in the same sentence, not a range.