Qwen3.8 27B compressed ~7x. Street llama.cpp + DFlash2 quote on one Spark.
Params
27B
Fastest decode
53 tok/s
DGX Spark · 2-bit class · llama.cpp
Refusal
Official
Setups
1
Best reported pair · not live serving speed ↗
Runs on
Quoted one-stream decode on real boxes.
No public weights link on this desk yet.