# Compare inference prices

Human: https://owterminal.com/inference/pricing
Agent: https://owterminal.com/inference/pricing.md

Reference snapshot: 2026-09-24. Source: https://openrouter.ai/api/v1/models.
USD per million tokens. These are reference quotes, not current pool offers.

| Family | Reference input | Reference output |
| --- | ---: | ---: |
| Qwen3.8 27B | $0.42 | $3 |
| Qwen3.8 Flash | $0.15 | $0.47 |
| Qwen3.6 35B-A3B | $0.15 | $1 |
| GLM 5.3 Flash | $0.15 | $0.5 |
| GLM 5.3 | $1.4 | $4.4 |
| GLM 5.2 | $0.65 | $2.04 |
| DeepSeek v4.1 Flash | $0.15 | $0.6 |
| Gemma 4 26B | $0.09 | $0.3 |
| Gemma 4 31B | $0.09 | $0.34 |
| Nemotron 3.5 Lightning | $0.08 | $0.2 |
| Ling 3.0 Flash | $0.021 | $0.063 |

Live offers appear on the human page and https://owterminal.com/api/v1/models. A host's one pool rate applies to both input and output tokens. Reported speed is not a guarantee. Different builds are not quality-equivalent. No automatic uncensored price surcharge applies.

## Supported builds

### GLM 5.3 Flash · EXL3

Alias: GLM-5.3-Flash-EXL3. Build: EXL3. Memory: Distributed / high-memory setup; use the model card's sizing. Runtime: Your existing compatible vLLM / EXL3 recipe.

License: Custom license · gated download; review terms before hosting. Source: https://huggingface.co/neko-legends/GLM-5.3-Flash-Uncensored-EXL3/tree/07135ec082f8f11f7a71e4244a4e5167a0f96277

Contribution multiplier: 1.5x. Host: https://owterminal.com/inference?door=host&model=GLM-5.3-Flash-EXL3

### Dolphin 3.0 · Llama 3.1 8B

Alias: Dolphin3.0-Llama3.1-8B. Build: BF16 source weights. Memory: ~16 GB weights only; add runtime and context-cache headroom. Runtime: vLLM or another compatible OpenAI endpoint.

License: Llama 3.1 Community License; review terms before hosting. Source: https://huggingface.co/cognitivecomputations/Dolphin3.0-Llama3.1-8B/tree/f065677950dfc7e708d518d64cf1f5041ee007a0

Contribution multiplier: 1.5x. Host: https://owterminal.com/inference?door=host&model=Dolphin3.0-Llama3.1-8B

### Qwen3.5 · 4B

Alias: Qwen3.5-4B. Build: BF16 source weights. Memory: ~8 GB weights only; add runtime and context-cache headroom. Runtime: Compatible vLLM / Transformers server.

License: Apache 2.0. Source: https://huggingface.co/Qwen/Qwen3.5-4B/tree/851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a

Contribution multiplier: 1x. Host: https://owterminal.com/inference?door=host&model=Qwen3.5-4B

## Contribution rules

Points are non-redeemable reputation, not credits, earnings or Lode. Both host and caller earn 1,000 points per $1 of eligible paid usage, with the model multiplier and a 1,000-point cap per account, role and UTC day. New settled ZEC/USDC deposits count; legacy balances, grants, failed jobs and self-calls do not. Distinct linked accounts, a reviewed declaration and an endpoint test within 24 hours are required. Pending period: 24 hours. Abuse awards may be reversed. Endpoint checks do not attest weights.

Account: https://owterminal.com/account/rewards
Use: https://owterminal.com/inference?door=call
Research: https://owterminal.com/models
