Decode · tok/s32Prefill—tok/100W13.3在单台 DGX Spark 上用 TensorFold 跑 Qwen3.8-Flash-Next,试着把 Vontra MLX 4-bit模型换成 NVIDIA NVFP4。结果 decode 仅仅32 tok/s。First-person swap on one Spark. Rank is the quoted 32 decode. Weight file 123.57 GiB.Officialclaimed@openaiarkaWatchCopy linkOpen pair