Choosing Embedding Models for Semantic Search in 2026

Choosing embedding models for semantic search in production comes down to five axes: retrieval quality on a benchmark that matches your languages, context length, dimension and its index cost, Matryoshka (MRL) support for dimension flexibility, and total operating cost across API calls or self-hosted GPUs. Managed APIs such as Cohere Embed v4 win on time-to-first-index; open weights such as Qwen3-Embedding win on cost at scale and data control. On the dimension axis, Cohere Embed v4 offers configurable output dimensions of 256, 512, 1024 or 1536, which already spans most vector-index budgets you will encounter.

Why the Model Choice Dominates

The embedding model sets the ceiling on retrieval quality: everything downstream — vector store, reranker, chunking — can only preserve or degrade what the encoder captured. Two behavioral properties deserve early attention. First, instruction-aware models change score when you change the prompt: Qwen3-Embedding accepts a task instruction on the query side, and omitting it degrades retrieval by roughly 1% to 5% according to the model card, so the instruction belongs in your embedding pipeline tests, not in a README. Second, asymmetric input types matter for commercial APIs. Cohere recommends embedding corpus passages with input_type="search_document" and user queries with input_type="search_query", because the prefixes were present at training time. Skipping them silently costs recall, and nothing in your dashboards will tell you why.

Compare the 2026 Contenders

The table below compares widely deployed models on the properties that decide production fit. Multilingual MTEB figures are the ones published by the model vendors themselves.

ModelWeightsContextDimensionsMRLNotes
Cohere Embed v4Closed API128k256–1536Yes, four stepsMultimodal text, images, mixed PDFs
Qwen3-Embedding-8BOpen, Apache-2.032kup to 4096YesInstruction-aware, 100+ languages
Qwen3-Embedding-0.6BOpen32kup to 1024YesCheap baseline for high-QPS query paths
multilingual-e5-large-instructOpen5121024NoLegacy but still deployed everywhere

Two structural differences drive the decision. Cohere Embed v4 handles text, images and interleaved PDF pages in a single model with a 128k-token window, which simplifies document-understanding workloads. Qwen3-Embedding-8B goes further on dimension control and multilingual reach: it scored 70.58 on the MTEB multilingual leaderboard, ranked No.1 as of June 5, 2025, a shortlisting signal rather than a procurement decision. For EU teams the licensing split also matters — Apache-2.0 weights can run inside your own perimeter, while the closed API routes corpus text through the vendor’s region.

Dimension, Context and Cost Math

Dimension is the biggest index-cost lever you control. A 4096-dim float32 vector costs about 16 KB per row; at 1024 dims it is 4 KB, and int8 quantization cuts that again by four. On 50 million chunks, moving from full-width float32 to 1024-dim int8 shrinks the raw vector payload from roughly 800 GB to about 50 GB — a change that shows up directly in the vector database bill and in p95 latency. Matryoshka-trained models make this tunable after the fact: Qwen3-Embedding-8B supports user-defined output dimensions from 32 to 4096, so you can embed once at full width and truncate per index without retraining. The pragmatic pattern is to embed at full dimension, build a first-stage index at a truncated cutoff, and A/B the cutoff against recall@50 before committing. Context length shapes chunking: 32k (Qwen3) comfortably covers RAG chunks of 512–2048 tokens plus metadata, while 128k (Cohere v4) lets you embed whole document pages. Serving is the other half of cost: Qwen3 ships with vLLM and Text Embeddings Inference recipes, so the batching and cache discipline from our KV cache and batching throughput guide applies to embedding servers too.

Run Your Own Benchmark

Leaderboard averages hide exactly the failures that hurt in production: your domain vocabulary, your languages, your chunk length. A one-day evaluation on your own traffic beats any MTEB mean:

  1. Collect 300–1000 real query/document pairs with judged relevance, even binary labels.
  2. Freeze chunking and the reranker so only the encoder varies.
  3. Embed the corpus with each candidate at full dimension and at one truncated MRL step.
  4. Measure nDCG@10, recall@50, and cost per million tokens (API price or GPU-hour equivalent).
  5. Select on recall per euro, then validate multilingual slices separately if you serve EU users.

For EU deployments, also verify where vectors are computed and stored: managed endpoints process your corpus in their region, while open weights keep the entire embedding step inside your perimeter — the same residency trade-off we detail in the multi-region AI API failover setup.

Deployment Recommendations

Default to these starting points. Prototype with a managed API to validate the product before owning GPUs. At sustained volumes above tens of millions of embeddings, move corpus indexing to self-hosted Qwen3-Embedding (4B or 8B) behind Text Embeddings Inference, and keep a 0.6B model on the latency-critical query path if QPS is high. Never mix models in one index: changing the encoder requires a full re-embed, so store the model ID and dimension alongside every vector and treat both as part of the index schema. Finally, budget the operational edge cases before going live — cold container starts on serverless embedding endpoints behave like the patterns covered in our serverless inference guide, and a query path that adds 900 ms of model-load latency will fail your SLOs even with perfect recall.

Sources