OpenWeightsTerminal
Pairs

Qwen3.8-Flash-Next

DGX Spark · NVFP4 · TensorFold

Decode · tok/s
32
Prefill
—
tok/100W
13.3
在单台 DGX Spark 上用 TensorFold 跑 Qwen3.8-Flash-Next,试着把 Vontra MLX 4-bit模型换成 NVIDIA NVFP4。结果 decode 仅仅32 tok/s。

First-person swap on one Spark. Rank is the quoted 32 decode. Weight file 123.57 GiB.

Officialclaimed
Open pair