Open weights, or rented thought.
Mega-corp inference is a tap. They open it. They meter it. They close it. Open weights are the only way a completion happens on silicon you can name — today, on a desk, without a landlord.

The lock is the product
Closed models do not publish a pair. There is no “their flagship on a 5090, 351 tok/s, claimed.” You cannot rank what you cannot run. That is not a missing benchmark. That is the product. If the weights stay closed, every token you generate sits on someone else’s policy, meter, and kill-switch. Inference becomes rent. Rent always has a landlord.
Do it now
The boxes are already here. A 12 GB 3060 is printing ~50 tok/s on a 35B-class distill. A 1660 SUPER is holding ~40. Flash-Next on an M5 Max peaked 126.5. Waiting for a cluster is how the window closes. Run the weight. File the pair. The street does not wait for a press kit.
Do it early
Spark desks are already splitting GLM as orchestrator and Qwen as worker. That is a stack, not a slide. Early means the day the GGUF lands, not the quarter a hyperscaler blesses an endpoint. Early is how open weights stay a market instead of a museum. File while the print is hot.
What this desk measures
Atomic unit: model × quant × runtime × chip. Rank is ordinal tok/s inside a filter. There is no composite score. Facts need a quote. Official, uncensored, and abliterated builds sit on the same ladder because the weights moved. That movement is the point — not a vibe, a pair.
The ladder
Tokens per 100 W
Hours to one million tokens
Quoted pairs vs unpublished
Where it ran
we <3 open weights. OWterminal is a product of dvidia.org. Idle silicon should serve. Weights should move. Completions should not have a landlord.