Tape
Ingest files chip, fit, tok/s, get URL, and whether someone is selling the box.
5090 bake: 2B at 351 tok/s, abliterated 26B-A4B at 173.
5060 Ti 16GB llama.cpp bake-off: 23 / 19 / 14 / 10 tok/s.
Flash-Next on M5 Max via MTPLX: 126.5 peak, 50 at 200k.
Official FP8 Flash-Next on 2× Spark, ~45 tok/s sustained.
Qwen3.8-27B UD-Q4_K_XL on 4090: 40.7 decode, 260k-class.
Abliterated 35B-A3B NVFP4 on 2× Spark: 94.4 decode.
Ornith ~40–46 tok/s on a 1660 SUPER 6GB.
Empero 35B-A3B Q4_K_M on a 12GB 3060: ~50 decode.
Locked suite on one M4 Pro: 85.5 tok/s at 21 GB peak.