# OWterminal > Measured tok/s for open and uncensored weights on real silicon. > Atomic unit: model × quant × runtime × chip. Rank is ordinal tok/s inside a filter. > There is no composite score. Site: https://owterminal.com Cite (HTML mirror): https://owterminal.com/cite Full dump: https://owterminal.com/llms-full.txt Languages: en, es, zh-Hans ## Pages - https://owterminal.com/ — Boards (ranked pairs) - https://owterminal.com/compare — Side-by-side bars - https://owterminal.com/models — Model index (NEW / HOT tags) - https://owterminal.com/hardware — Chip index - https://owterminal.com/market — Used-box street tape - https://owterminal.com/get — Hugging Face / get links - https://owterminal.com/tape — Source posts - https://owterminal.com/runtimes — llama.cpp, MLX, MTPLX - https://owterminal.com/cite — Citation desk for humans and models - https://owterminal.com/llms.txt — this file - https://owterminal.com/llms-full.txt — every quoted pair ## How to cite Quote the pair, not a brand ranking. Example: "Qwen3.5 2B Q4_K_M on RTX 5090 32GB, 351 tok/s decode, claimed by @fillagrew." Status is claimed | reproduced | harness-locked. Refusal is official | uncensored | abliterated | heretic. ## Models (12) - Qwen3.8-27B — Qwen, dense, 27B, official [NEW, HOT] https://huggingface.co/Qwen/Qwen3.8-27B - Qwen3.6-35B-A3B — Qwen, MoE 35B-A3B, 35B / 3B act, official https://huggingface.co/Qwen/Qwen3.6-35B-A3B - Qwen3.6-35B-A3B abliterated — community, MoE 35B-A3B, 35B / 3B act, abliterated [HOT] - Empero Qwen3.8-35B-A3B — empero-ai, MoE distill, 35B / 3B act, official [NEW, HOT] - Qwen3.8-Flash-Next — Qwen, dense, unknown, official [NEW, HOT] - Qwen3.5 2B — Qwen, dense, 2B, official [NEW, HOT] - Qwen3.5-4B — Qwen, dense, 4B, official [NEW] - Qwen3.5-9B — Qwen, dense, 9B, official [NEW] - Qwen3-8B — Qwen, dense, 8B, official - Gemma-4 26B-A4B-it — Google / community, MoE, 26B-A4B, abliterated [NEW, HOT] - Nex-N2.5-mini — community, MoE post-train, 35B-A3B class, official [NEW] - Ornith — ornith-ai, unknown, unknown, official [NEW] https://huggingface.co/ornith-ai ## Chips (8) - M4 Pro — Apple, unified, 4 pairs - M5 Max — Apple, unified (SKU unknown), 1 pairs - DGX Spark — NVIDIA, 2× Spark, 2 pairs - RTX 5090 32GB — NVIDIA, 32 GB VRAM, 2 pairs - RTX 4090 24GB — NVIDIA, 24 GB VRAM, 1 pairs - RTX 5060 Ti 16GB — NVIDIA, 16 GB VRAM, 4 pairs - RTX 3060 12GB — NVIDIA, 12 GB VRAM, 1 pairs - GTX 1660 SUPER 6GB — NVIDIA, 6 GB VRAM + 40 GB RAM, 1 pairs ## Fastest pairs (decode tok/s) 1. Qwen3.5 2B on RTX 5090 32GB — 351 tok/s (Q4_K_M GGUF, unknown, official) 2. Gemma-4 26B-A4B-it on RTX 5090 32GB — 173 tok/s (Q4_K_M GGUF, unknown, abliterated) 3. Qwen3.8-Flash-Next on M5 Max — 126.5 tok/s (unknown, MTPLX V2.11.3, official) 4. Qwen3.6-35B-A3B on 2× DGX Spark — 94.4 tok/s (NVFP4, unknown, abliterated) 5. Qwen3.6-35B-A3B on M4 Pro — 85.5 tok/s (4-bit MLX, rapid-mlx, official) 6. Qwen3.5-4B on M4 Pro — 82.8 tok/s (4-bit MLX, rapid-mlx, official) 7. Empero Qwen3.8-35B-A3B on RTX 3060 12GB — 50 tok/s (Q4_K_M GGUF, llama.cpp, official) 8. Qwen3.5-9B on M4 Pro — 49.3 tok/s (4-bit MLX, rapid-mlx, official) 9. Qwen3-8B on M4 Pro — 48.3 tok/s (4-bit MLX, rapid-mlx, official) 10. Qwen3.8-Flash-Next on 2× DGX Spark — 45 tok/s (FP8, unknown, official) 11. Qwen3.8-27B on RTX 4090 24GB — 40.7 tok/s (UD-Q4_K_XL GGUF, llama.cpp, official) 12. Ornith on GTX 1660 SUPER 6GB — 40 tok/s (unlinked, unknown, official) 13. Qwen3.8-27B on RTX 5060 Ti 16GB — 23 tok/s (GSQ-RCO IQ3_XXS-mtp, llama.cpp, official) 14. Qwen3.8-27B on RTX 5060 Ti 16GB — 19 tok/s (UD-Q2_K_XL, llama.cpp, official) 15. Nex-N2.5-mini on RTX 5060 Ti 16GB — 14 tok/s (Q4_K_M GGUF, llama.cpp, official) 16. Qwen3.8-27B TurboFCFusion on RTX 5060 Ti 16GB — 10 tok/s (IQ2_M GGUF, llama.cpp, uncensored) ## Tape - 2026-09-17 @fillagrew — 5090 bake: 2B at 351 tok/s, abliterated 26B-A4B at 173. https://x.com/fillagrew/status/2100478211134025937 - 2026-09-17 @fntAInhead — 5060 Ti 16GB llama.cpp bake-off: 23 / 19 / 14 / 10 tok/s. https://x.com/fntAInhead/status/2100499089431359808 - 2026-09-17 @Youssofal_ — Flash-Next on M5 Max via MTPLX: 126.5 peak, 50 at 200k. https://x.com/Youssofal_/status/2100468205030719533 - 2026-09-17 @Oluwaphilemon1 — Official FP8 Flash-Next on 2× Spark, ~45 tok/s sustained. https://x.com/Oluwaphilemon1/status/2100498960645308573 - 2026-09-17 @Oluwaphilemon1 — Qwen3.8-27B UD-Q4_K_XL on 4090: 40.7 decode, 260k-class. https://x.com/Oluwaphilemon1/status/2100409183115821394 - 2026-09-17 @bonellisystems — Abliterated 35B-A3B NVFP4 on 2× Spark: 94.4 decode. https://x.com/bonellisystems/status/2100417191434768450 - 2026-09-17 @mine_craft_bui — Ornith ~40–46 tok/s on a 1660 SUPER 6GB. https://x.com/mine_craft_bui/status/2100492671534166056 - 2026-09-17 @Oluwaphilemon1 — Empero 35B-A3B Q4_K_M on a 12GB 3060: ~50 decode. https://x.com/Oluwaphilemon1/status/2100374395482939401 - 2026-09-16 @rapidmlx — Locked suite on one M4 Pro: 85.5 tok/s at 21 GB peak. https://x.com/rapidmlx/status/2100252655990010209