mofcloud

Embedding Model Price Comparison

What does it do?
Turns text into vectors — for knowledge-base retrieval, semantic search and recommendations.
How's it priced?
By input tokens only; there's no output price.
What are dimensions?
The vector length. Bigger is finer but heavier on storage; 1024–1536 is most common.
Is open source free?
Weights are free and self-hostable — you just bring your own GPU.
Monthly cost?
How much you vectorize × the unit price. For volume, estimate with the Token calculator.
Prices updated 2026-08-28
Model Maker Price / 1M tokens Dimensions Max input Tags
BGE-M3 智源 BAAI 🇨🇳 Open · free 1024 8K Chinese
Qwen3-Embedding-8B 阿里通义 🇨🇳 Open · free 4096 32K Chinese
text-embedding-3-small OpenAI 🇺🇸 $0.020 1536 8K
voyage-3.5-lite Voyage AI 🇺🇸 $0.020 1024 32K
jina-embeddings-v3 Jina AI 🇩🇪 $0.020 1024 8K
voyage-3.5 Voyage AI 🇺🇸 $0.060 1024 32K
text-embedding-v4 阿里通义 🇨🇳 $0.074 2048 8K Chinese
text-embedding-ada-002 OpenAI 🇺🇸 $0.100 1536 8K
Mistral Embed Mistral 🇫🇷 $0.100 1024 8K
embed v4.0 Cohere 🇨🇦 $0.120 1536 128K Multimodal
text-embedding-3-large OpenAI 🇺🇸 $0.130 3072 8K
gemini-embedding-001 Google 🇺🇸 $0.150 3072 2K
voyage-3-large Voyage AI 🇺🇸 $0.180 1024 32K
voyage-code-3 Voyage AI 🇺🇸 $0.180 1024 32K Code

Prices are indicative — verify on each vendor's official page. International models are priced in USD (¥ converted at the latest rate); Chinese models use official CNY. Compiled from vendor pricing pages and models.dev.

Guide

How to choose an embedding model?

Embeddings are usually called at far higher volume than chat — the whole knowledge base gets vectorized once, then every query is encoded too, so a small unit-price gap balloons into a big monthly bill. If budget matters pick a cheap one; to save the most, self-host open BGE / Qwen3-Embedding.

Beyond price, watch two numbers: dimensions set how fine the vector is and how much your vector DB costs to store and search (1024–1536 is plenty; go 3072 for accuracy); max input caps how much fits per call — most sit near 8K so long docs need chunking, while a few like Cohere embed v4 reach 128K.

Chinese models (Tongyi, BGE) bill in CNY and handle Chinese well; for multilingual or long context, Cohere, Voyage and Jina are common. To turn this into a monthly number, plug your volume into the Token cost calculator.