mofcloud

Pick an LLM by use case — which is cheapest?

Data updated 2026-08-29

Typical workload: ~80,000 input tokens · ~1,500 output tokens per call · context ≥ 128K

All qualifying models (156)
# Model In / Out / 1M Context Per call vs cheapest
1 Qwen Flash 🇨🇳 Alibaba Cloud $0.02 / $0.22 1000K $0.0021
2 Qwen3.7 Flash 🇨🇳 Alibaba Cloud $0.03 / $0.12 1000K $0.0025 ×1.2
3 Qwen3.5 Flash 🇨🇳 Alibaba Cloud $0.03 / $0.30 1000K $0.0028 ×1.4
4 Qwen Turbo 🇨🇳 Alibaba Cloud $0.04 / $0.09 1000K $0.0037 ×1.8
5 GPT-5 Nano 🇺🇸 OpenAI $0.05 / $0.40 400K $0.0046 ×2.2
6 Qwen2.5 7B Instruct 🇨🇳 Alibaba Cloud $0.07 / $0.14 131K $0.0060 ×2.9
7 Qwen Long 🇨🇳 Alibaba Cloud $0.07 / $0.29 10000K $0.0062 ×3.0
8 GLM-4.7-FlashX 🇨🇳 Zhipu GLM $0.07 / $0.40 200K $0.0062 ×3.0
9 GLM-5.3-Flash 🇨🇳 Zhipu GLM $0.07 / $0.25 1000K $0.0064 ×3.1
10 Qwen Doc Turbo 🇨🇳 Alibaba Cloud $0.09 / $0.14 131K $0.0072 ×3.4
11 Mistral Small 3.2 🇫🇷 Mistral $0.10 / $0.30 128K $0.0085 ×4.1
12 Gemini 2.5 Flash-Lite 🇺🇸 Google $0.10 / $0.40 1049K $0.0086 ×4.1

Need an exact figure from your real prompt? Open the Token Cost Calculator →

How to use

Why does “the cheapest” change from one scenario to the next?

Most people pick a model off a single “price per million tokens” list, but the real bill depends on how long your input and output actually are. On the same price sheet, long-doc summary (80k input, short output) and complex reasoning (short input, huge output) surface almost entirely different “cheapest” models.

This tool presets a typical token workload for each scenario, plugs it into every model's official input and output prices, and ranks per-call cost. So you don't fill in numbers — just pick a scenario.

What each pick means
  • 💰 Best value: the lowest per-call cost that meets the scenario's hard requirements (context, modality, capability) — for high-volume and batch work. Usually a Chinese model in this data.
  • 🏁 Top Chinese: the strongest qualifying Chinese flagship (DeepSeek / Qwen / GLM / Kimi, etc.) — strong Chinese, compliance-friendly, for when quality matters.
  • 🌐 Top global: the strongest qualifying Western flagship (GPT / Claude / Gemini) — when the business mandates a Western vendor or you want the best.

Want it exact? Each scenario uses a “typical” workload; yours may differ. Paste your real text into the Token Cost Calculator for a precise figure, or see every spec in the LLM price comparison.

FAQ

How is the ranking produced?

Each scenario presets a typical input/output token workload plus hard requirements (context window, modality, capabilities), plugs them into each model's official input and output prices, and sorts by per-call cost ascending. No subjective scoring.

Why do different scenarios recommend different models?

Because cost depends on your input and output lengths. Input-heavy scenarios (long docs, RAG) are driven by input price; output-heavy ones (reasoning, code) by output price. On the same price sheet, the cheapest often differs between them.

How are the Chinese and global flagships chosen?

Flagships come from a hand-curated list of each major lab's top-tier model (DeepSeek V4 Pro, Qwen Max, GPT-5.5 Pro, Claude Opus, etc.). Top Chinese is the highest-ranked Chinese flagship that meets the scenario; Top global likewise for Western labs. They represent the “don't compromise on capability” choice and usually cost far more than the value pick.

Are these prices accurate?

Prices reuse the main pricing dataset from vendors' official docs — Chinese models in official CNY, global models converted at the latest rate. For reference; vendor billing is authoritative.

Can I compute from my own prompt?

Yes. This uses typical per-scenario values; paste your real text into the Token Cost Calculator to compute per-call and monthly cost from your actual input/output lengths and call volume.

---