# OWterminal > Measured tok/s for open and uncensored weights on real silicon. > Atomic unit: model × quant × runtime × chip. Rank is ordinal tok/s inside a filter. > There is no composite score. Site: https://owterminal.com Cite (HTML mirror): https://owterminal.com/cite Full dump: https://owterminal.com/llms-full.txt Languages: en, es, zh-Hans ## Pages - https://owterminal.com/ — Boards (ranked pairs) - https://owterminal.com/compare — Side-by-side bars - https://owterminal.com/models — Model index (NEW / HOT tags) - https://owterminal.com/hardware — Chip index - https://owterminal.com/market — Used-box street tape - https://owterminal.com/get — Hugging Face / get links - https://owterminal.com/tape — Source posts - https://owterminal.com/runtimes — llama.cpp, MLX, MTPLX - https://owterminal.com/cite — Citation desk for humans and models - https://owterminal.com/talk — DGXtalk, human floor (no model posts) - https://owterminal.com/labs — Labs, nations, humanity points, GDP model (labeled, not measured) - https://owterminal.com/watts — tok/s per 100W of named TDP - https://owterminal.com/manifesto — Open weights, or rented thought - https://owterminal.com/terms — Terms of service - https://owterminal.com/agreement — Desk agreement - https://owterminal.com/pool — Bring a machine, or fund a key - https://owterminal.com/letter — The letter - https://owterminal.com/llms.txt — this file - https://owterminal.com/llm.txt — same file, the short path models ask for - https://owterminal.com/llms-full.txt — every quoted pair ## Model pages - https://owterminal.com/model/qwen3.8-27b — Qwen3.8-27B (Qwen) - https://owterminal.com/model/qwen3.6-35b-a3b — Qwen3.6-35B-A3B (Qwen) - https://owterminal.com/model/qwen3.6-35b-a3b-abliterated — Qwen3.6-35B-A3B abliterated (community) - https://owterminal.com/model/empero-qwen3.8-35b-a3b — Empero Qwen3.8-35B-A3B (empero-ai) - https://owterminal.com/model/qwen3.8-flash-next — Qwen3.8-Flash-Next (Qwen) - https://owterminal.com/model/qwen3.5-2b — Qwen3.5 2B (Qwen) - https://owterminal.com/model/qwen3.5-4b — Qwen3.5-4B (Qwen) - https://owterminal.com/model/qwen3.5-9b — Qwen3.5-9B (Qwen) - https://owterminal.com/model/qwen3-8b — Qwen3-8B (Qwen) - https://owterminal.com/model/gemma-4-26b-a4b-it-abliterated — Gemma-4 26B-A4B-it (Google / community) - https://owterminal.com/model/nex-n2.5-mini — Nex-N2.5-mini (community) - https://owterminal.com/model/ornith — Ornith (ornith-ai) - https://owterminal.com/model/glm-5.3-flash — GLM 5.3 Flash (Zhipu) - https://owterminal.com/model/deepseek-v4.1-flash — DeepSeek v4.1 Flash (DeepSeek) - https://owterminal.com/model/mimo-v2.6-distill-qwen-9b — MiMo-V2.6-Distill-Qwen-9B (Xiaomi MiMo) - https://owterminal.com/model/aliceai-80b-a3b — AliceAI-Foundation-80B-A3B (Yandex) ## Chip pages - https://owterminal.com/chip/m4-pro — M4 Pro (Apple) - https://owterminal.com/chip/m5-max — M5 Max (Apple) - https://owterminal.com/chip/dgx-spark — DGX Spark (NVIDIA) - https://owterminal.com/chip/rtx-5090-32gb — RTX 5090 (NVIDIA) - https://owterminal.com/chip/rtx-4090-24gb — RTX 4090 (NVIDIA) - https://owterminal.com/chip/rtx-5060-ti-16gb — RTX 5060 Ti (NVIDIA) - https://owterminal.com/chip/rtx-5060-8gb — RTX 5060 (NVIDIA) - https://owterminal.com/chip/rtx-3060-12gb — RTX 3060 (NVIDIA) - https://owterminal.com/chip/gtx-1660-super-6gb — GTX 1660 SUPER (NVIDIA) ## How to cite Quote the pair, not a brand ranking. Example: "Qwen3.5 2B Q4_K_M on RTX 5090 32GB, 351 tok/s decode, claimed by @fillagrew." Do not invent tok/s. Closed APIs have no pair until they publish one. Status is claimed | reproduced | harness-locked. Refusal is official | uncensored | abliterated | heretic. ## Models (16) - Qwen3.8-27B — Qwen, dense, 27B, official [NEW, HOT] https://huggingface.co/Qwen/Qwen3.8-27B - Qwen3.6-35B-A3B — Qwen, MoE 35B-A3B, 35B / 3B act, official https://huggingface.co/Qwen/Qwen3.6-35B-A3B - Qwen3.6-35B-A3B abliterated — community, MoE 35B-A3B, 35B / 3B act, abliterated [HOT] - Empero Qwen3.8-35B-A3B — empero-ai, MoE distill, 35B / 3B act, official [NEW, HOT] - Qwen3.8-Flash-Next — Qwen, dense, unknown, official [NEW, HOT] - Qwen3.5 2B — Qwen, dense, 2B, official [NEW, HOT] - Qwen3.5-4B — Qwen, dense, 4B, official [NEW] - Qwen3.5-9B — Qwen, dense, 9B, official [NEW] - Qwen3-8B — Qwen, dense, 8B, official - Gemma-4 26B-A4B-it — Google / community, MoE, 26B-A4B, abliterated [NEW, HOT] - Nex-N2.5-mini — community, MoE post-train, 35B-A3B class, official [NEW] - Ornith — ornith-ai, unknown, unknown, official [NEW] https://huggingface.co/ornith-ai - GLM 5.3 Flash — Zhipu, dense, unknown, official [NEW, HOT] - DeepSeek v4.1 Flash — DeepSeek, dense, unknown, official [NEW, HOT] - MiMo-V2.6-Distill-Qwen-9B — Xiaomi MiMo, dense, 9B, official [NEW, HOT] - AliceAI-Foundation-80B-A3B — Yandex, MoE 80B-A3B, 80B / 3B act, official [NEW] https://huggingface.co/Yamada114514/AliceAI-Foundation-80B-A3B-Base-GGUF ## Chips (9) - M4 Pro — Apple, 24–64 GB unified, 4 pairs - M5 Max — Apple, 64–128 GB unified, 2 pairs - DGX Spark — NVIDIA, 128 GB unified, 3 pairs - RTX 5090 — NVIDIA, 32 GB VRAM, 3 pairs - RTX 4090 — NVIDIA, 24 GB VRAM, 1 pairs - RTX 5060 Ti — NVIDIA, 8–16 GB VRAM, 4 pairs - RTX 5060 — NVIDIA, 8 GB VRAM, 1 pairs - RTX 3060 — NVIDIA, 8–12 GB VRAM, 1 pairs - GTX 1660 SUPER — NVIDIA, 6 GB VRAM, 1 pairs ## Fastest pairs (decode tok/s) 1. Qwen3.5 2B on RTX 5090 32GB — 351 tok/s (Q4_K_M GGUF, unknown, official) https://owterminal.com/pair/qwen35-2b-q4km-5090 2. Gemma-4 26B-A4B-it on RTX 5090 32GB — 173 tok/s (Q4_K_M GGUF, unknown, abliterated) https://owterminal.com/pair/gemma4-26b-abliterated-q4km-5090 3. Qwen3.8-Flash-Next on M5 Max — 126.5 tok/s (unknown, MTPLX V2.11.3, official) https://owterminal.com/pair/qwen38-flash-mtplx-m5max 4. Qwen3.6-35B-A3B on 2× DGX Spark — 94.4 tok/s (NVFP4, unknown, abliterated) https://owterminal.com/pair/qwen36-35b-abliterated-nvfp4-spark 5. Qwen3.6-35B-A3B on M4 Pro — 85.5 tok/s (4-bit MLX, rapid-mlx, official) https://owterminal.com/pair/qwen36-35b-4bit-mlx-m4pro 6. Qwen3.5-4B on M4 Pro — 82.8 tok/s (4-bit MLX, rapid-mlx, official) https://owterminal.com/pair/qwen35-4b-4bit-mlx-m4pro 7. Qwen3.8-Flash-Next on RTX 5090 32GB — 80.9 tok/s (NVFP4, unknown, official) https://owterminal.com/pair/qwen38-flash-nvfp4-5090 8. Qwen3.8-Flash-Next on DGX Spark — 79.5 tok/s (EXL3 3.05 bpw, EXL3, official) https://owterminal.com/pair/qwen38-flash-exl3-spark 9. AliceAI-Foundation-80B-A3B on M5 Max 128GB — 65 tok/s (Q4_K_M GGUF, llama.cpp, official) https://owterminal.com/pair/aliceai-80b-q4-m5max 10. Empero Qwen3.8-35B-A3B on RTX 3060 12GB — 50 tok/s (Q4_K_M GGUF, llama.cpp, official) https://owterminal.com/pair/empero-qwen38-35b-q4km-3060 11. Qwen3.5-9B on M4 Pro — 49.3 tok/s (4-bit MLX, rapid-mlx, official) https://owterminal.com/pair/qwen35-9b-4bit-mlx-m4pro 12. Qwen3-8B on M4 Pro — 48.3 tok/s (4-bit MLX, rapid-mlx, official) https://owterminal.com/pair/qwen3-8b-4bit-mlx-m4pro 13. MiMo-V2.6-Distill-Qwen-9B on RTX 5060 8GB — 47 tok/s (Q5_K_M GGUF, llama.cpp, official) https://owterminal.com/pair/mimo-v26-qwen9b-q5-5060 14. Qwen3.8-Flash-Next on 2× DGX Spark — 45 tok/s (FP8, unknown, official) https://owterminal.com/pair/qwen38-flash-fp8-spark 15. Qwen3.8-27B on RTX 4090 24GB — 40.7 tok/s (UD-Q4_K_XL GGUF, llama.cpp, official) https://owterminal.com/pair/qwen38-27b-ud-q4k-xl-4090 16. Ornith on GTX 1660 SUPER 6GB — 40 tok/s (unlinked, unknown, official) https://owterminal.com/pair/ornith-1660-super 17. Qwen3.8-27B on RTX 5060 Ti 16GB — 23 tok/s (GSQ-RCO IQ3_XXS-mtp, llama.cpp, official) https://owterminal.com/pair/qwen38-27b-gsq-rco-5060ti 18. Qwen3.8-27B on RTX 5060 Ti 16GB — 19 tok/s (UD-Q2_K_XL, llama.cpp, official) https://owterminal.com/pair/qwen38-27b-ud-q2k-xl-5060ti 19. Nex-N2.5-mini on RTX 5060 Ti 16GB — 14 tok/s (Q4_K_M GGUF, llama.cpp, official) https://owterminal.com/pair/nex-n25-mini-q4km-5060ti 20. Qwen3.8-27B TurboFCFusion on RTX 5060 Ti 16GB — 10 tok/s (IQ2_M GGUF, llama.cpp, uncensored) https://owterminal.com/pair/qwen38-27b-turbofc-uncen-5060ti ## Tape - 2026-09-22 @stfu0911 — MiMo-V2.6 Distill Qwen-9B Q5_K_M on RTX 5060 8GB: ~47 decode, ~1600 prefill, 262k. https://x.com/stfu0911/status/2102409155143409688 - 2026-09-22 @aqty — AliceAI 80B-A3B Q4_K_M on M5 Max 128GB: ~65 tok/s, 45.1 GiB. https://x.com/aqty/status/2102413529206976799 - 2026-09-21 @tekizaihq — Flash-Next NVFP4 on one 5090: 80.9 tok/s single-stream. https://x.com/tekizaihq/status/2101844907618939233 - 2026-09-20 @yume_arasaki — Flash-Next EXL3 on one Spark: 79.5 code reproduced; 102.6 repetitive clamps; 71.7 prose. https://x.com/yume_arasaki/status/2101741448811229219 - 2026-09-20 @ViC305 — EXL3 Spark recipe reproduced at 79.5 on Yume’s box. https://x.com/ViC305/status/2101747348322103706 - 2026-09-20 @MiaAI_lab — Spark street: Flash-Next solo; GLM 5.3 / DeepSeek v4.1 Flash on 2×+; orch+worker+scanner at 6×. https://x.com/MiaAI_lab/status/2101821681404624976 - 2026-09-21 @sethforprivacy — 8 Sparks: 4× GLM ring, 2× Flash-Next, 2× DeepSeek v4f. https://x.com/sethforprivacy/status/2101834433087070708 - 2026-09-17 @fillagrew — 5090 bake: 2B at 351 tok/s, abliterated 26B-A4B at 173. https://x.com/fillagrew/status/2100478211134025937 - 2026-09-17 @fntAInhead — 5060 Ti 16GB llama.cpp bake-off: 23 / 19 / 14 / 10 tok/s. https://x.com/fntAInhead/status/2100499089431359808 - 2026-09-17 @Youssofal_ — Flash-Next on M5 Max via MTPLX: 126.5 peak, 50 at 200k. https://x.com/Youssofal_/status/2100468205030719533 - 2026-09-17 @Oluwaphilemon1 — Official FP8 Flash-Next on 2× Spark, ~45 tok/s sustained. https://x.com/Oluwaphilemon1/status/2100498960645308573 - 2026-09-17 @Oluwaphilemon1 — Qwen3.8-27B UD-Q4_K_XL on 4090: 40.7 decode, 260k-class. https://x.com/Oluwaphilemon1/status/2100409183115821394 - 2026-09-17 @bonellisystems — Abliterated 35B-A3B NVFP4 on 2× Spark: 94.4 decode. https://x.com/bonellisystems/status/2100417191434768450 - 2026-09-17 @mine_craft_bui — Ornith ~40–46 tok/s on a 1660 SUPER 6GB. https://x.com/mine_craft_bui/status/2100492671534166056 - 2026-09-17 @Oluwaphilemon1 — Empero 35B-A3B Q4_K_M on a 12GB 3060: ~50 decode. https://x.com/Oluwaphilemon1/status/2100374395482939401 - 2026-09-16 @rapidmlx — Locked suite on one M4 Pro: 85.5 tok/s at 21 GB peak. https://x.com/rapidmlx/status/2100252655990010209 ## Street stacks (DGX Spark) Recipes and roles, not tok/s. Source: https://x.com/MiaAI_lab/status/2101821681404624976 - 1× Solo: qwen3.8-flash-next (Most common solo load) https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark - 2× Dual: glm-5.3-flash (Default dual) https://github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks; deepseek-v4.1-flash https://github.com/MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks; qwen3.8-flash-next https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks - 3× Triple: glm-5.3-flash + qwen3.8-flash-next (GLM orch on 2×, Qwen worker on 1×); glm-5.3-flash (Homogeneous — Mia’s own desk); deepseek-v4.1-flash (Homogeneous) - 4× Quad: glm-5.3-flash + qwen3.8-flash-next (GLM orch 2× + Qwen worker 2×); glm-5.3-flash (Homogeneous); deepseek-v4.1-flash (Homogeneous) - 6× 6+: glm-5.3-flash + qwen3.8-flash-next + deepseek-v4.1-flash (GLM orch + Qwen worker + DeepSeek scanner)