Claude Sonnet 5.5
It solves the most real coding tasks in independent tests.
Choose AI
GeneralIntelligence ranking for everything in generalCodingCoding tools and agentsCreate imagesImage generatorsCreate videosVideo generatorsFree and on your computerOpen source, private and unlimitedChoose AI / Coding
Which AI best solves real coding tasks (installing, debugging, compiling…), where to use it and how much it costs.
Three ways to get it right: the best one regardless of price, the best value and the best one without paying.
It solves the most real coding tasks in independent tests.
It scores 52 versus 64 for the most powerful one, and it costs about 15 times less to use.
Score with its «high» reasoning mode, the best value; in its most powerful mode it reaches 56.
It's open source: you download it and use it without paying and without usage limits. It scores 35 versus 64 for the most powerful one.
Needs a computer with about 625 GB of video memory (VRAM).
One model per row (its best version). The variants (max, high, medium…) are in «For the data nerds», at the end.
For developers: models that don't write text, but pick an option, give a score or find the most similar item. They're used to filter, label or match.
Measured with the Decision Index (real decision tasks, corrected for chance). The table has 17 decision models, 6 embedding models and 3 rerankers.
It's the best at deciding in the Decision Index.
⚠ Ojo: Barely tested: no independent ranking has confirmed it yet; Little used: almost nobody uses it yet. That's why we don't recommend it by default.
It scores 60 versus 63 for the most powerful one.
It's open source: you download it and use it without paying and without usage limits. It scores 57 versus 63 for the most powerful one.
Needs a computer with about 20.0 GB of video memory (VRAM).
26 models and variants · click ▸ on a row to see sources, memory and notes
| Model | Perf.DI / MTEB ⇅ | Price€/1M$/1M tok ⇅ | WeightsWeights⇅ | MaturityMad.⇅ | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
Perplexity | 62.8 | self-host | OpenOpen | Barely testedBarely t. ⬇ 599/mo HF · 1 rankingLittle usedLittle used | |||||||
Tipo: Typed decision Featured: ★ N.º 1 (open) Weights: OpenOpen 🤗 Hugging Face · ↗ Web Maturity: Barely testedBarely t. Very little use · 599 HF downloads/mo · 1 leaderboard (Decision Index 0.3) · 4 days · Adoption: Little used (⬇ 599/mo HF · 1 ranking) (leaderboards: Decision Index 0.3; released: 2026-10-05; data from 2026-10-09) Score: #1 in the Decision Index, above Jev. What it does: Typed decision (choice/score/yes-no with probabilities) on Qwen3.8-27B with the causal mask lifted and a 255-option decision head. Memory: disk 52.2 GB (BF16) · Min. VRAM 49.0 GB official (weights) · fits on: 1× 80 GB (BF16, official)
Requirement: Model card: «a CUDA GPU with room for roughly 49 GiB of weights plus working memory». source Quality: Decision Index 0.3: 62.75 (Full score, #1 of 115; Jev 60.11) · public index 62.25 · the card states 61.56 on 0.2.1 · Latency: Median 104.1 ms (Decision Index 0.3, 1× RTX PRO 6000). · Price: Self-host (your own GPU) v1 (= «AutoJev-27B», 56.4 on 0.2.1) → v1.1 (Oct 5): more data (tasksource) and no causal mask. Its own code (autojev); the quantizations are from the community. License: Apache 2.0 Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-08 | |||||||||||
| Fastino GLiDE (28B, no-thinking) Fastino | 60.2 | n/d | ClosedClosed | NewNew 1 DI ranking · weights not publicNo usage dataNo data | |||||||
Tipo: Typed decision Featured: ★ N.º 2 DI Weights: ClosedClosed ↗ Web Maturity: NewNew By default: in Decision Index 0.3.1 as #2 · weights not public yet · Adoption: No usage data (1 DI ranking · weights not public) (leaderboards: Decision Index 0.3.1; released: 2026-09-30; data from 2026-10-10) Score: #2 in the DI; weights «coming soon». What it does: «Thinking»/no-thinking decision model: choice, yes/no and score with probabilities; does not generate free text. Weights «coming soon» according to DI. Memory: Weights not public yet Quality: Decision Index 0.3.1: Full 60.21 (#2 of 116) · public 59.06. Only behind Perplexity Decider v1.1 (62.75). · Latency: Median 98.4 ms (Decision Index 0.3.1). · Price: n/a (API; public price not listed in DI) Entry from the official 0.3.1 leaderboard (generated 2026-10-10). DI metadata: «Open weights coming soon». Not to be confused with GLiNER. Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-10 | |||||||||||
| Jev 1.13 (System One) TypeSafe AI | 60.1 | €0.037$0.042/1M | ClosedClosed | NewNew OpenRouter top 20 · #9 · 1 rankingWidely usedWidely used | |||||||
Tipo: Typed decision Featured: ★ Best value (widely used) Weights: ClosedClosed ↗ Web Maturity: NewNew Released 22 days ago · top 20 OpenRouter · 1 leaderboard (Decision Index 0.3) · Adoption: Widely used (OpenRouter top 20 · #9 · 1 ranking) (leaderboards: Decision Index 0.3; released: 2026-09-17; data from 2026-10-09) Score: The only API in the top; 64k context; can't be fine-tuned. What it does: Reads a state (text/JSON) and answers typed questions: choice (up to 255 options), score (2–10 levels) or yes/no, with probabilities and confidence. Without generating text. Memory: Hosted (not downloadable) Quality: Decision Index 0.3: 60.11 (Full score, #3 of 115) · public index 57.96 · Latency: Median 524 ms per request measured by the Decision Index (hosted API, network round trip); TypeSafe advertises 70–500 ms. · Price: $0.042/1M input tokens; output free In Decision Index 0.3 (Oct 7) it comes 3rd, tied with Fastino GLiDE (weights «coming soon») and behind Perplexity Decider v1.1 (open). 64k context; best in English; cannot be fine-tuned. Also on Vercel AI Gateway (typesafe-ai/jev). Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-08 | |||||||||||
| Torchcast Decision 27B Torchcast AI | 59.9 | self-host | OpenOpen | Barely testedBarely t. ⬇ 602/mo HF · 1 rankingLittle usedLittle used | |||||||
Tipo: Typed decision Weights: OpenOpen 🤗 Hugging Face · ↗ Web Maturity: Barely testedBarely t. Very little use · 602 HF downloads/mo · 1 leaderboard (Decision Index 0.3) · 6 days · Adoption: Little used (⬇ 602/mo HF · 1 ranking) (leaderboards: Decision Index 0.3; released: 2026-10-03; data from 2026-10-09) Score: Best public index, but non-commercial license. What it does: Typed decision on Qwen3.8-27B (merged LoRA); 256k context. Memory: disk 54.7 GB (BF16) · Min. VRAM 20.0 GB approx. · fits on: 1× 24 GB (approx.)
Quality: Decision Index 0.3: 59.91 (Full, n.º 4) · public index 65.10 (#1) · the index guided the choice of checkpoint (declared) · Latency: Median 99.9 ms (Decision Index 0.3) · card: 33 ms per yes/no question on 1× H100. · Price: Self-host (your own GPU) Big public jump (65.1) that shrinks on the private tests (Full 59.9). There is also a 12B version. License: CC BY-NC 4.0 (non-commercial) Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-08 | |||||||||||
| Kev 27B Jared Palmer | 58.8 | self-host | OpenOpen | NewNew ⬇ 1.5k/mo HF · 1 rankingLittle usedLittle used | |||||||
Tipo: Typed decision Weights: OpenOpen 🤗 Hugging Face · ↗ Web Maturity: NewNew Released 15 days ago · 2k HF downloads/mo · 1 leaderboard (Decision Index 0.3) · Adoption: Little used (⬇ 1.5k/mo HF · 1 ranking) (leaderboards: Decision Index 0.3; released: 2026-09-24; data from 2026-10-09) Score: The best calibrated in the top; data with a commercial license. What it does: Typed decision (full fine-tune of Qwen3.8-27B + head); states of up to 64k tokens. Memory: disk 51.2 GB (BF16) · Min. VRAM 65.5 GB official · fits on: 1× 80 GB (H100/H200/B200, official)
Requirement: Model card: 51 GB weights, server ~65.5 GB resident; 64k states reach 87.1 GB (H200). source Quality: Decision Index 0.3: 58.78 (Full, n.º 6) · public 56.69 · best score on private new domains among open models (55.5) · ECE 0.022 · Latency: Median 135 ms (Decision Index 0.3). · Price: Self-host (your own GPU) Card: data from sources with licenses that allow commercial use, and contamination screening. English only. License: Apache 2.0 Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-08 | |||||||||||
| Blink v0.3 26B-A4B Pixilab.ai | 57.8 | self-host | OpenOpen | Barely testedBarely t. ⬇ 765/mo HF · 1 rankingLittle usedLittle used | |||||||
Tipo: Typed decision Featured: ★ Jev level at 25 ms Weights: OpenOpen 🤗 Hugging Face · ↗ Web Maturity: Barely testedBarely t. Very little use · 765 HF downloads/mo · 1 leaderboard (Decision Index 0.3) · 7 days · Adoption: Little used (⬇ 765/mo HF · 1 ranking) (leaderboards: Decision Index 0.3; released: 2026-10-02; data from 2026-10-09) Score: Jev level at 25 ms on 1 GPU. What it does: Typed decision on Gemma 4 26B-A4B (MoE, 4B active), published already quantized (NVFP4 / FP8). Memory: disk 18.8 GB (NVFP4) · Min. VRAM 23.0 GB approx. · fits on: 1× 32 GB (approx.)
Requirement: Model card: «17.5 GB on disk, served on a single RTX PRO 5000 with 32k context». source Quality: Decision Index 0.3: 57.76 (Full, n.º 8) · public 58.21 · ECE 0.040 · Latency: Median 25.1 ms (Decision Index 0.3): the fastest in the top 10. · Price: Self-host (your own GPU) NVFP4 for Blackwell GPUs; FP8 build for Hopper/Ada. Card: 32k context on an RTX PRO 5000. License: Apache 2.0 Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-08 | |||||||||||
| Rune 26B-A4B v3 Surogate | 57.4 | self-host | OpenOpen | NewNew ⬇ 5.2k/mo HF · 1 rankingModerately usedModerately used | |||||||
Tipo: Typed decision Weights: OpenOpen ↗ Web Maturity: NewNew Released 18 days ago · 5k HF downloads/mo · 1 leaderboard (Decision Index 0.3) · Adoption: Moderately used (⬇ 5.2k/mo HF · 1 ranking) (leaderboards: Decision Index 0.3; released: 2026-09-21; data from 2026-10-09) Score: MoE 4B active; community GGUF. What it does: Open typed-decision model (full fine-tune of Gemma 4 26B-A4B). Memory: disk 51.6 GB (n/d) · Min. VRAM 20.0 GB approx. · fits on: 1× 24 GB (approx.)
Quality: Decision Index 0.3: 57.43 (Full score, #10 of 115) · public index 58.28 · Latency: Median 120.5 ms (Decision Index 0.3, 1× RTX PRO 6000, BF16, sequential client). · Price: Self-host (your own GPU) Almost Jev's level, open; in BF16 it needs 1× 80 GB, quantized 1× 24–32 GB (approx.). License: Apache 2.0 (repo gated) Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-08 | |||||||||||
| Bespoke Nimble 9B v3 Bespoke Labs | 54.7 | self-host | OpenOpen | Barely testedBarely t. ⬇ 189/mo HF · 1 rankingLittle usedLittle used | |||||||
Tipo: Typed decision Weights: OpenOpen 🤗 Hugging Face · ↗ Web Maturity: Barely testedBarely t. Very little use · 189 HF downloads/mo · 1 leaderboard (Decision Index 0.3) · 9 days · Adoption: Little used (⬇ 189/mo HF · 1 ranking) (leaderboards: Decision Index 0.3; released: 2026-09-30; data from 2026-10-09) Score: Best ≤10B; non-commercial license. What it does: Typed-decision LoRA adapter for Qwen3.5-9B (choice or yes/no with a probability per option). Memory: disk 17.9 GB (LoRA 0.69 GB + base) · Min. VRAM 9.0 GB approx. · fits on: 1× 8–12 GB (approx.)
Disk = merged BF16 GGUF; the official adapter is 0.69 GB and needs the Qwen3.5-9B base. Quality: Decision Index 0.3: 54.67 (Full, #18; best ≤10B) · public 57.19 · Latency: Median 73.9 ms (Decision Index 0.3). · Price: Self-host (your own GPU) The official repo is only the LoRA (0.69 GB) on Qwen3.5-9B (Apache 2.0); ggml-org publishes merged GGUFs. License: CC BY-NC 4.0 (non-commercial) Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-08 | |||||||||||
| Drex v1.5 Nace.AI | 50.8 | €0.036$0.040/1M | OpenOpen | NewNew HF ~98 downloads/mo · OpenRouter · 1 rankingLittle usedLittle used | |||||||
Tipo: Typed decision Featured: ★ Open ≤10B (public) Weights: OpenOpen 🤗 Hugging Face · ↗ Web Maturity: NewNew By default: released ~1 day ago · in Decision Index 0.3.1 · open weights + API · Adoption: Little used (HF ~98 downloads/mo · OpenRouter · 1 ranking) (leaderboards: Decision Index 0.3.1; released: 2026-10-09; data from 2026-10-10) Score: Open ≤10B with the best public index; Full drops to 50.8. RAIL-M license. What it does: Typed decisions (choice / noul yes-no / score) over a text/JSON state; one pass per question, without generating text. Compatible with Jev's System One (/v1/systemone). Memory: disk 17.9 GB (BF16) · Min. VRAM 24.0 GB official · fits on: 1× GPU 24 GB (A10G bf16, test Nace) · Q8_0 ~9.5 GB on Metal/CPU
Requirement: Model card: bf16 ~18 GB on a CUDA GPU (tested on A10G 24 GB); Q8_0 GGUF ~9.5 GB on Apple silicon or CPU. source Q8_0 isn't available as a ready GGUF repo: it's converted with Nace's llama.cpp fork (drex-v1.5 branch). Q8 GB according to the model card. Quality: Decision Index 0.3.1: Full 50.84 (#24 of 116) · public index 58.08 (best open ≤10B on public; 0.9-pt tie with Jev/Nimble on public). JevBench 86.2% (Jev 87.0%). · Latency: Median 53.5 ms (Decision Index 0.3.1, official board). OpenRouter P50 ~0.18 s (DeepInfra). Model card: ~0.65 s (8k–32k) · ~2.0 s (32k–128k) on long docs. · Price: Self-host · OpenRouter API $0.04/1M input (output free) Weights on HF (2026-10-09 on OpenRouter). Base XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B. 16k context by default, up to 131k. Own API drex.nace.ai and OpenRouter nace-ai/drex-v1.5 at $0.04/1M in. In DI Full it falls far behind Decider/GLiDE/Jev on private tests (weak on knowledge: e.g. GPQA). Its own forks of llama.cpp/Ollama. License: Nace.AI Open RAIL-M (restricted use; not Apache) Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-10 | |||||||||||
| ezjev 4B s2 everettjf | 47.0 | self-host | OpenOpen | Barely testedBarely t. ⬇ 83/mo HF · 1 rankingLittle usedLittle used | |||||||
Tipo: Typed decision Weights: OpenOpen 🤗 Hugging Face · ↗ Web Maturity: Barely testedBarely t. Very little use · 83 HF downloads/mo · 1 leaderboard (Decision Index 0.3) · 6 days · Adoption: Little used (⬇ 83/mo HF · 1 ranking) (leaderboards: Decision Index 0.3; released: 2026-10-03; data from 2026-10-09) Score: Best ~4B; check licenses before commercial use. What it does: Typed decision on Qwen3.5-4B (merged LoRA), served with vLLM. Memory: disk 9.1 GB (BF16) · Min. VRAM 5.0 GB approx. · fits on: 1× 8–12 GB (approx.)
Quality: Decision Index 0.3: 46.95 (Full, #35; best ~4B) · public 50.82 · trained with train/dev splits of ~17 benchmarks from the index (declared) · Latency: Median 35.5 ms (Decision Index 0.3, 1× RTX PRO 6000). · Price: Self-host (your own GPU) The author now recommends ezjev-4b-s3 (no score on the board). Some training sources are non-commercial (e.g. ANLI). License: Other (MIT code; data with NC sources) Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-08 | |||||||||||
| Decider 4B Mapika | 40.3 | self-host | OpenOpen | NewNew ⬇ 25k/mo HF · 1 rankingModerately usedModerately used | |||||||
Tipo: Typed decision Weights: OpenOpen ↗ Web Maturity: NewNew Released 17 days ago · 25k HF downloads/mo · 1 leaderboard (Decision Index 0.3) · Adoption: Moderately used (⬇ 25k/mo HF · 1 ranking) (leaderboards: Decision Index 0.3; released: 2026-09-22; data from 2026-10-09) Score: 13 ms; open recipe (Mapika). What it does: Typed decisions with calibrated probabilities in one pass (decider family 2B/4B/35B-A3B). Memory: disk 8.4 GB (n/d) · Min. VRAM 11.0 GB approx. · fits on: 1× 16 GB (approx.)
Quality: Decision Index 0.3: 40.33 (Full score, #50 of 115) · public index 41.19 · Latency: Median 12.6 ms (Decision Index 0.3, FP8 + CUDA graphs, 1× RTX PRO 6000). · Price: Self-host (your own GPU) Decider family 2B / 4B / 35B-A3B (35B-A3B: Full 48.5). Permissive license; recipe at github.com/Mapika/decider. License: Apache 2.0 Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-08 | |||||||||||
| Liquid AI d1-3B Liquid AI | 40.1 | self-host | OpenOpen | NewNew ⬇ 5.4k/mo HF · 1 rankingModerately usedModerately used | |||||||
Tipo: Typed decision Featured: ★ Best edge Weights: OpenOpen ↗ Web Maturity: NewNew Released 4 days ago · 5k HF downloads/mo · 1 leaderboard (Decision Index 0.3) · Adoption: Moderately used (⬇ 5.4k/mo HF · 1 ranking) (leaderboards: Decision Index 0.3; released: 2026-10-05; data from 2026-10-09) Score: Edge: 6 GB, 16 ms; text + image. What it does: Decision model: a single pass, zero output tokens; choice/score/yes-no with confidence and probabilities. Text + image. Memory: disk 6.2 GB (BF16) · Min. VRAM 4.0 GB approx. · fits on: Any 8 GB GPU, Jetson or Mac
Requirement: Model card: 8 ms per decision on an RTX 4090 (bf16); also Jetson Orin Nano (50 ms) and Apple M5 Pro (30 ms). source Quality: Decision Index 0.3: 40.08 (Full score, #51 of 115) · public index 48.99 · Latency: Median 16.1 ms (Decision Index 0.3; Liquid AI submission on AMD MI325X) · model card: 8 ms on an RTX 4090, 30 ms Apple M5 Pro, 50 ms Jetson Orin Nano. · Price: Self-host (your own GPU) It scores 49.0 on the public index, but drops on the private tests (Full 40.1). 32k context. License: LFM Open License 1.0 Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-08 | |||||||||||
| Liquid AI d1-omni-600M Liquid AI | 9.4 | self-host | OpenOpen | NewNew ⬇ 8.5k/mo HF · 1 rankingModerately usedModerately used | |||||||
Tipo: Typed decision Weights: OpenOpen ↗ Web Maturity: NewNew Released 4 days ago · 8k HF downloads/mo · 1 leaderboard (Decision Index 0.3) · Adoption: Moderately used (⬇ 8.5k/mo HF · 1 ranking) (leaderboards: Decision Index 0.3; released: 2026-10-05; data from 2026-10-09) Score: Text, image and audio; for edge. What it does: Mini version of d1: decisions on text, image and audio with zero output tokens. Memory: disk 2.4 GB (n/d) · Min. VRAM 2.0 GB approx. · fits on: 1× 8–12 GB (approx.) Quality: Decision Index 0.3: 9.45 (Full score, #97 of 115) · public index 17.90 · Latency: Median 10.9 ms (Decision Index 0.3, Liquid AI submission). · Price: Self-host (CPU/edge) For cheap mass filtering or on-device; 16k context. License: LFM Open License 1.0 Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-08 | |||||||||||
| Laya Convai Innovations | 4.4 | self-host | OpenOpen | NewNew ⬇ 36k/mo HF · 1 rankingWidely usedWidely used | |||||||
Tipo: Typed decision Weights: OpenOpen 🤗 Hugging Face · ↗ Web Maturity: NewNew Released 21 days ago · 36k HF downloads/mo · 1 leaderboard (Decision Index 0.3) · Adoption: Widely used (⬇ 36k/mo HF · 1 ranking) (leaderboards: Decision Index 0.3; released: 2026-09-18; data from 2026-10-09) Score: Super fast, but ~chance without fine-tuning: only useful fine-tuned with your labels. What it does: Non-autoregressive encoder (ModernBERT-large + decision head, trained with RLCD): choice/score/yes-no in one pass. Family: laya (English, 421M), laya-multilingual (mmBERT-base, 322M) and laya-typed-decisions (fine-tuned). Memory: disk 0.8 GB (F16) · Min. VRAM 2.0 GB approx. · fits on: Any GPU (CPU possible, slower)
Quality: Decision Index 0.3: 4.43 (Full, #106 of 115) · public 6.19 · 0.2.1: 6.04 · card: 0.36 zero-shot accuracy (chance 0.32); 0.766 only the checkpoint fine-tuned on that same task · Latency: Median 5.8 ms (Decision Index 0.3, CUDA graphs, 1× RTX PRO 6000) · the model card states 33 ms per question on a T4 GPU. · Price: Self-host · free on Vercel AI Gateway until Oct 31 (promo, @vercel_dev Oct 1) The card's «better than Jev» comparisons use other samples and a checkpoint fine-tuned on that task; Jev was not measured on the same suite. Limited options (≲20). Servable with llama.cpp (/v1/systemone, ggml-org/Laya-GGUF). Hosted: free on Vercel AI Gateway until Oct 31 (https://x.com/vercel_dev/status/2105758650052849859); also on Dell Enterprise Hub. License: Apache 2.0 Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-08 | |||||||||||
Alibaba | MTEB-M 70.6 / 64.3 | self-host | OpenOpen | ProvenProven ⬇ 11.2M/mo HF · 1 rankingWidely usedWidely used | |||||||
Tipo: Embeddings · retrieve Featured: ★ Retrieve Weights: OpenOpen ↗ Web Maturity: ProvenProven Massive use · 11.2 M HF downloads/mo · 1 leaderboard (MTEB) · from 2025-06 · Adoption: Widely used (⬇ 11.2M/mo HF · 1 ranking) (leaderboards: MTEB; released: 2025-06-03; data from 2026-10-09) Score: Multilingual, 32k, configurable dimension. What it does: Multilingual embeddings (100+ languages) to retrieve the most similar texts by similarity; configurable dimension (MRL) and per-task instructions. Memory: disk 15.1 GB (n/d) · Min. VRAM 3.0 GB approx. · fits on: 0.6B: CPU or any GPU · 8B: 1× 24 GB (approx.) Quality: MTEB multilingual: 70.58 (8B) / 64.33 (0.6B) · Latency: n/d · Price: Self-host (your own GPU) 32k context (long documents). The 0.6B is the sweet spot for indexing millions of texts. License: Apache 2.0 Source: huggingface.co/Qwen/Qwen3-Embedding-8B · 2026-10-08 | |||||||||||
Google | Best <500M on MTEB (Google) | self-host | OpenOpen | ProvenProven ⬇ 3.6M/mo HF · 1 rankingWidely usedWidely used | |||||||
Tipo: Embeddings · retrieve Weights: OpenOpen ↗ Web Maturity: ProvenProven Massive use · 3.6 M HF downloads/mo · 1 leaderboard (MTEB) · from 2025-07 · Adoption: Widely used (⬇ 3.6M/mo HF · 1 ranking) (leaderboards: MTEB; released: 2025-07-17; data from 2026-10-09) Score: On-device; 2k context (split long documents). What it does: Very light multilingual embeddings (308M) for devices or CPU; 768→128 dimensions (MRL). Memory: disk 1.2 GB (n/d) · Min. VRAM 0.2 GB official (RAM, QAT) · fits on: CPU / mobile (<200 MB of RAM quantized, QAT)
Requirement: Google: runs in less than 200 MB of RAM with quantization (QAT). source Quality: Best open multilingual embedding <500M on MTEB (according to Google) · Latency: <15 ms with 256 tokens (EdgeTPU) · Price: Self-host (CPU) 2k context: short for long documents (split them). License: Gemma Source: developers.googleblog.com/en/introducing-embeddinggemma/ · 2026-10-08 | |||||||||||
| BGE-M3 BAAI | MTEB-M 59.6 | self-host | OpenOpen | ProvenProven ⬇ 33.2M/mo HF · 1 rankingWidely usedWidely used | |||||||
Tipo: Embeddings · retrieve Weights: OpenOpen ↗ Web Maturity: ProvenProven Massive use · 33.2 M HF downloads/mo · 1 leaderboard (MTEB) · from 2024-01 · Adoption: Widely used (⬇ 33.2M/mo HF · 1 ranking) (leaderboards: MTEB; released: 2024-01-27; data from 2026-10-09) Score: Dense + lexical + ColBERT; 8k. What it does: Hybrid embedding: dense + lexical (sparse) + multi-vector (ColBERT) in one model; useful when exact terms matter (names, codes). Memory: disk 2.3 GB (n/d) · Min. VRAM 4.0 GB approx. · fits on: 1× 8–12 GB (approx.)
Quality: Multilingual MTEB: 59.56 (Qwen table, May 2025) · Latency: n/d · Price: Self-host (your own GPU or CPU) 8192 context. Veteran and well tested; lower score than Qwen3 but native hybrid search. License: MIT Source: huggingface.co/BAAI/bge-m3 · 2026-10-08 | |||||||||||
| Voyage voyage-4 / 4-large Voyage AI | n/d | €0.054$0.060/1M | ClosedClosed | ProvenProven 1 rankingNo usage dataNo data | |||||||
Tipo: Embeddings · retrieve Weights: ClosedClosed ↗ Web Maturity: ProvenProven Veteran · 1 leaderboard (MTEB) · from 2026-01 · Adoption: No usage data (1 ranking) (leaderboards: MTEB; released: 2026-01-15; data from 2026-10-09) Score: 200M tokens free; lite $0.02/1M. What it does: Embeddings via API with no infrastructure. Memory: Hosted Quality: n/d · Latency: n/d · Price: $0.06/1M (voyage-4) · $0.12/1M (4-large); 200M tokens free voyage-4-lite $0.02/1M. Batch API −33 %. Source: docs.voyageai.com/docs/pricing · 2026-10-08 | |||||||||||
Alibaba | MMTEB-R 72.7 / 66.4 | self-host | OpenOpen | ProvenProven ⬇ 3.9M/mo HF · 0 rankingsWidely usedWidely used | |||||||
Tipo: Rerankers · rerank Featured: ★ Rerank Weights: OpenOpen ↗ Web Maturity: ProvenProven Massive use · 3.9 M HF downloads/mo · since 2025-05 · no independent leaderboard in the sources consulted · Adoption: Widely used (⬇ 3.9M/mo HF · 0 rankings) (released: 2025-05-29; data from 2026-10-09) Score: Cross-encoder with instructions; 32k. What it does: Cross-encoder: scores each (query, document) pair by reading both at once; reranks the top-k returned by the embedding. Memory: disk 8.0 GB (n/d) · Min. VRAM 3.0 GB approx. · fits on: 0.6B: any GPU · 4B: 1× 12–16 GB (approx.) Quality: MMTEB-R: 72.74 (4B) / 66.36 (0.6B); BGE-reranker-v2-m3: 58.36 · Latency: n/d · Price: Self-host (your own GPU) 32k context; accepts an instruction («does this text answer the question?»). License: Apache 2.0 Source: huggingface.co/Qwen/Qwen3-Reranker-0.6B · 2026-10-08 | |||||||||||
| bge-reranker-v2-m3 BAAI | MMTEB-R 58.4 | self-host | OpenOpen | ProvenProven ⬇ 16.6M/mo HF · 0 rankingsWidely usedWidely used | |||||||
Tipo: Rerankers · rerank Weights: OpenOpen ↗ Web Maturity: ProvenProven Massive use · 16.6 M HF downloads/mo · since 2024-03 · no independent leaderboard in the sources consulted · Adoption: Widely used (⬇ 16.6M/mo HF · 0 rankings) (released: 2024-03-15; data from 2026-10-09) Score: Partner of BGE-M3. jina-reranker-v3 is non-commercial. What it does: Lightweight multilingual cross-encoder; natural partner of BGE-M3. Memory: disk 2.3 GB (n/d) · Min. VRAM 4.0 GB approx. · fits on: 1× 8–12 GB (approx.)
Quality: MMTEB-R: 58.36 (Qwen table) · Latency: n/d · Price: Self-host (your own GPU or CPU) Alternative: jina-reranker-v3 (0.6B) is CC-BY-NC, not suitable for commercial use. License: Apache 2.0 Source: huggingface.co/BAAI/bge-reranker-v2-m3 · 2026-10-08 | |||||||||||
| Voyage rerank-3 / rerank-3-lite Voyage AI | n/d | €0.045$0.050/1M | ClosedClosed | Barely testedBarely t. 0 rankingsNo usage dataNo data | |||||||
Tipo: Rerankers · rerank Weights: ClosedClosed ↗ Web Maturity: Barely testedBarely t. No independent benchmark · 38 days · Adoption: No usage data (0 rankings) (released: 2026-09-01; data from 2026-10-09) Score: ~$0.0025 per query of 100 docs. What it does: Reranker via API. Memory: Hosted Quality: n/d · Latency: n/d · Price: $0.05/1M tokens (rerank-3) · $0.02/1M (lite); ~$0.0025 per request of 100 docs Billed tokens = query tokens × no. of docs + doc tokens. 200M free. Source: docs.voyageai.com/docs/pricing · 2026-10-08 | |||||||||||
OpenAI | n/a (not in the DI) | €0.089$0.10/1M | ClosedClosed | Barely testedBarely t. 0 rankingsNo usage dataNo data | |||||||
Tipo: Typed decision Weights: ClosedClosed ↗ Web Maturity: Barely testedBarely t. No independent benchmark · 3 days · Adoption: No usage data (0 rankings) (released: 2026-10-06; data from 2026-10-09) Score: Direct Jev rival (public beta Oct 6); accepts images. What it does: /v1/decisions endpoint: predicate (yes/no), choice or score questions over text or images, with a structured answer; uses GPT-6 Luna. Memory: Hosted Quality: Not in Decision Index 0.3. · Latency: OpenAI: up to 10× faster than GPT-6 Luna via the Responses API. No independent measurement. · Price: $0.10/1M input tokens; output and cache free Jev costs $0.042/1M input and only accepts text. OpenAI acknowledges that Jev «inspired» the product. Source: fortune.com/2026/10/08/jev-an-ai-for-making-quick-decisions-has-been-a · 2026-10-09 | |||||||||||
Microsoft | n/a (not in DI yet) | €0.037$0.042/1M | ClosedClosed | NewNew OpenRouter (new) · 0 rankings DINo usage dataNo data | |||||||
Tipo: Typed decision Featured: ★ New MS API Weights: ClosedClosed ↗ Web Maturity: NewNew By default: released 1 day ago · no Decision Index yet · API on OpenRouter · Adoption: No usage data (OpenRouter (new) · 0 rankings DI) Score: Same price as Jev ($0.042/1M in); weights not public. What it does: Typed decisions (yes/no, choice, score/rubric) with calibrated probabilities; post-trained on Qwen3.5-9B. Decisions API (Foundry + OpenRouter), does not generate text. Memory: Hosted Quality: Not in Decision Index 0.3.1 (snapshot 2026-10-09). Microsoft: best accuracy on its pack of 36 blind benchmarks (~150k questions) versus H2O-Lightning-4B, Rune, Quyet, GPT-6 Luna Decisions, etc. · Latency: OpenRouter P50 ~0.30 s (Azure). Microsoft: P50 ~35× faster than GPT-6 Sol on its internal bench. · Price: $0.042/1M input; output free (same as Jev) Launched 2026-10-09 on Foundry and OpenRouter. Closed weights (open base Qwen3.5-9B). Xbox Research: competitive with GPT-6 Sol at >14× the speed and ~200× lower cost in labelling. Ideal for matching/routing/classification. Source: commandline.microsoft.com/microsoft-decision-1-model-foundry/ · 2026-10-10 | |||||||||||
| JEV-27B-VL (AutoTrust) AutoTrust AI | DI vision 69.8 (#1) | self-host | OpenOpen | NewNew ⬇ 1.5M/mo HF · 1 rankingWidely usedWidely used | |||||||
Tipo: Typed decision Weights: OpenOpen 🤗 Hugging Face · ↗ Web Maturity: NewNew Released 9 days ago · 1.5 M HF downloads/mo · 1 leaderboard (Decision Index (vision)) · Adoption: Widely used (⬇ 1.5M/mo HF · 1 ranking) (leaderboards: Decision Index (vision); released: 2026-09-30; data from 2026-10-09) Score: The best in the Decision Index with images. What it does: Typed decision over images + text (LoRA adapter on a 27B base); /v1/decide endpoint in vLLM. Memory: disk 55.6 GB (BF16) · Min. VRAM 2.0 GB approx. · fits on: 1× 8–12 GB (approx.)
AutoTrust says there are quantized variants; not verified here. Quality: Decision Index vision board: 69.82 Full (#1), 72.78 public; 2nd Solomon v1.1 with 67.66. Its sibling GEV-26B-Decide gets 56.17 on the general 0.3 index. · Latency: Median 263.9 ms per row on the Decision Index vision board. · Price: Self-host (your own GPU) License: Apache 2.0 (adapters); base under its own license Source: huggingface.co/spaces/multimodalart/jev-decision-index · 2026-10-09 | |||||||||||
Google | MTEB Code 78.7 | self-host | OpenOpen | Barely testedBarely t. ⬇ 21k/mo HF · 0 rankingsWidely usedWidely used | |||||||
Tipo: Embeddings · retrieve Weights: OpenOpen 🤗 Hugging Face · ↗ Web Maturity: Barely testedBarely t. No independent benchmark · 21k HF downloads/mo · 25 days · Adoption: Widely used (⬇ 21k/mo HF · 0 rankings) (released: 2026-09-14; data from 2026-10-09) Score: Successor to the 300m: Apache 2.0, 8k context (4×). What it does: Multimodal embeddings (text, code, image, video and audio) in a single space; 768→128 dimensions (MRL); 8k context. Memory: disk 1.5 GB (BF16) · Min. VRAM 0.6 GB official (RAM, quantized) · fits on: CPU / mobile (~191 MB of RAM text only, ~567 MB multimodal, quantized)
Requirement: Google: ~191 MB of active RAM text only and ~567 MB multimodal, quantized, on a Pixel 11 Pro. source Quality: Google: <1B multimodal leader on MTEB Code and MAEB; MTEB Code 68.76 → 78.68 versus EmbeddingGemma 1. · Latency: On-device; Google gives no latency. · Price: Self-host (CPU/mobile) License: Apache 2.0 Source: blog.google/innovation-and-ai/technology/developers-tools/embeddinggem · 2026-10-09 | |||||||||||
Perplexity | ViDoRe v3 62.3 / 65.2 | self-host | OpenOpen | NewNew ⬇ 1k/mo HF · 1 rankingLittle usedLittle used | |||||||
Tipo: Embeddings · retrieve Weights: OpenOpen 🤗 Hugging Face · ↗ Web Maturity: NewNew Only 1 leaderboard (MTEB) · 1k HF downloads/mo · 81 days · Adoption: Little used (⬇ 1k/mo HF · 1 ranking) (leaderboards: MTEB; released: 2026-07-20; data from 2026-10-09) Score: Multi-vector: more accurate, but the index takes up more space. What it does: Multi-vector retrieval (ColBERT, MaxSim) of text, images and PDFs without OCR; both sizes share the space (index with 9B, query with 0.6B). Memory: disk 2.4 GB (F32) · Min. VRAM 5.0 GB approx. · fits on: 1× 8–12 GB (approx.) Weights published in F32; in BF16 they'd be half. Quality: ViDoRe v3 nDCG@10 (image): 62.3 (0.6B) and 65.2 (9B), according to the model cards. · Latency: n/d · Price: Self-host (your own GPU) License: MIT Source: aicoder.com/news/news-20261008-perplexity-pplx-embed-v2-late-multimoda · 2026-10-09 | |||||||||||
Decision Index 0.3 (official board multimodalart/jev-decision-index, data from Oct 7): «Full score» = 20% public benchmarks + 50% private tests of the same skills + 30% private tasks from new domains, corrected for chance (0 = chance, 100 = perfect). Latency = median per request measured by the board (1× RTX PRO 6000, sequential client; Jev = API over the network). MTEB/MMTEB-R: tables from the Qwen model cards.
Quality vs price chart, full table with filters and all the variants of each model (max, high, medium…). Click ▸ on a row to see sources, memory and notes.
Sources publish API prices in dollars. With the €/$ button at the top you choose the currency; in euros they are converted at €1 = $1.1206 (ECB reference rate of October 9, 2026).
Up = better · left = cheaper · the dashed line is the efficient frontier.
Swipe to see the whole chart
No cost or no data: Gemma 4 31B, Mellum2.1.
40 models and variants · click ▸ on a row to see sources, memory and notes
| Model | TB 4.0% solved ⇅ | Costper task ⇅ | WeightsWeights⇅ | MaturityMad.⇅ | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
Anthropic | 63.6 | €11.29$12.65 | ClosedClosed | NewNew on OpenRouter · 5 rankingsModerately usedModerately used | |||||||
Featured: on the efficient frontier (nothing gives more performance for less money) Weights: ClosedClosed Maturity: NewNew Released 11 days ago · on OpenRouter · 5 leaderboards · Adoption: Moderately used (on OpenRouter · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-28; data from 2026-10-09) AA Coding Agents Index: 68.4 with “Claude Code - Sonnet 5.5 (max)” at €12.66$14.19/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Anthropic | 59.6 | €11.70$13.11 | ClosedClosed | ProvenProven OpenRouter top 20 · #2 · 5 rankingsWidely usedWidely used | |||||||
Weights: ClosedClosed Maturity: ProvenProven top 20 OpenRouter · 5 leaderboards · 17 days · Adoption: Widely used (OpenRouter top 20 · #2 · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-22; data from 2026-10-09) AA Coding Agents Index: 66.0 with “Claude Code - Opus 5.5 (max)” at €11.63$13.04/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Anthropic | 59.6 | €7.83$8.78 | ClosedClosed | ProvenProven OpenRouter top 20 · #2 · 5 rankingsWidely usedWidely used | |||||||
Weights: ClosedClosed Maturity: ProvenProven top 20 OpenRouter · 5 leaderboards · 17 days · Adoption: Widely used (OpenRouter top 20 · #2 · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-22; data from 2026-10-09) Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
OpenAI | 59.6 | €5.23$5.86 | ClosedClosed | ProvenProven on OpenRouter · 5 rankingsModerately usedModerately used | |||||||
Featured: on the efficient frontier (nothing gives more performance for less money) Weights: ClosedClosed Maturity: ProvenProven on OpenRouter · 5 leaderboards · 36 days · Adoption: Moderately used (on OpenRouter · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-03; data from 2026-10-09) Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
OpenAI | 59.1 | €7.58$8.50 | ClosedClosed | ProvenProven on OpenRouter · 5 rankingsModerately usedModerately used | |||||||
Weights: ClosedClosed Maturity: ProvenProven on OpenRouter · 5 leaderboards · 36 days · Adoption: Moderately used (on OpenRouter · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-03; data from 2026-10-09) AA Coding Agents Index: 61.6 with “Codex - GPT-6 Astra (max)” at €6.66$7.47/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Anthropic | 57.1 | €5.42$6.08 | ClosedClosed | NewNew on OpenRouter · 5 rankingsModerately usedModerately used | |||||||
Weights: ClosedClosed Maturity: NewNew Released 11 days ago · on OpenRouter · 5 leaderboards · Adoption: Moderately used (on OpenRouter · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-28; data from 2026-10-09) AA Coding Agents Index: 62.9 with “Claude Code - Sonnet 5.5 (xhigh)” at €2.97$3.33/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Google | 57.1 | €6.05$6.78 | ClosedClosed | NewNew 5 rankingsNo usage dataNo data | |||||||
Weights: ClosedClosed Maturity: NewNew Released 9 days ago · 5 leaderboards · Adoption: No usage data (5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-30; data from 2026-10-09) Score: In the agents index it is listed as unavailable. AA Coding Agents Index: 63.8 with “Antigravity CLI - Gemini 4 Argon (high)” at €5.21$5.84/task (marked as unavailable) Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Anthropic | 56.6 | €4.57$5.12 | ClosedClosed | ProvenProven OpenRouter top 20 · #2 · 5 rankingsWidely usedWidely used | |||||||
Featured: on the efficient frontier (nothing gives more performance for less money) Weights: ClosedClosed Maturity: ProvenProven top 20 OpenRouter · 5 leaderboards · 17 days · Adoption: Widely used (OpenRouter top 20 · #2 · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-22; data from 2026-10-09) Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
OpenAI | 56.1 | €1.63$1.82 | ClosedClosed | ProvenProven OpenRouter top 20 · #1 · 5 rankingsWidely usedWidely used | |||||||
Featured: on the efficient frontier (nothing gives more performance for less money) Weights: ClosedClosed Maturity: ProvenProven top 20 OpenRouter · 5 leaderboards · 10 days · Adoption: Widely used (OpenRouter top 20 · #1 · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-29; data from 2026-10-09) AA Coding Agents Index: 60.1 with “Codex - GPT-6.1 Sol (max)” at €1.39$1.55/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Anthropic | 55.1 | €14.08$15.78 | ClosedClosed | ProvenProven on OpenRouter · 5 rankingsModerately usedModerately used | |||||||
Weights: ClosedClosed Maturity: ProvenProven on OpenRouter · 5 leaderboards · 38 days · Adoption: Moderately used (on OpenRouter · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-01; data from 2026-10-09) Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-09 | |||||||||||
OpenAI | 54.0 | €0.92$1.03 | ClosedClosed | ProvenProven OpenRouter top 20 · #1 · 5 rankingsWidely usedWidely used | |||||||
Featured: on the efficient frontier (nothing gives more performance for less money) Weights: ClosedClosed Maturity: ProvenProven top 20 OpenRouter · 5 leaderboards · 10 days · Adoption: Widely used (OpenRouter top 20 · #1 · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-29; data from 2026-10-09) AA Coding Agents Index: 62.9 with “Codex - GPT-6.1 Sol (xhigh)” at €0.93$1.04/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Anthropic | 52.0 | €17.15$19.22 | ClosedClosed | ProvenProven on OpenRouter · 5 rankingsModerately usedModerately used | |||||||
Weights: ClosedClosed Maturity: ProvenProven on OpenRouter · 5 leaderboards · 38 days · Adoption: Moderately used (on OpenRouter · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-01; data from 2026-10-09) AA Coding Agents Index: 62.2 with “Claude Code - Fable 5.1 (max) (with fallback)” at €11.05$12.39/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-09 | |||||||||||
OpenAI | 51.5 | €0.74$0.83 | ClosedClosed | ProvenProven OpenRouter top 20 · #1 · 5 rankingsWidely usedWidely used | |||||||
Featured: ★ Best value · on the efficient frontier (nothing gives more performance for less money) Weights: ClosedClosed Maturity: ProvenProven top 20 OpenRouter · 5 leaderboards · 10 days · Adoption: Widely used (OpenRouter top 20 · #1 · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-29; data from 2026-10-09) AA Coding Agents Index: 60.1 with “Codex - GPT-6.1 Sol (high)” at €0.79$0.89/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
OpenAI | 48.0 | €0.55$0.61 | ClosedClosed | ProvenProven OpenRouter top 20 · #1 · 5 rankingsWidely usedWidely used | |||||||
Featured: on the efficient frontier (nothing gives more performance for less money) Weights: ClosedClosed Maturity: ProvenProven top 20 OpenRouter · 5 leaderboards · 10 days · Adoption: Widely used (OpenRouter top 20 · #1 · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-29; data from 2026-10-09) AA Coding Agents Index: 61.4 with “Codex - GPT-6.1 Sol (medium)” at €0.63$0.70/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Z AI | 41.9 | €7.20$8.06 | OpenOpen | ProvenProven OpenRouter top 20 · #4 · 5 rankingsWidely usedWidely used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 1.5 M HF downloads/mo · top 20 OpenRouter · 5 leaderboards · 52 days · Adoption: Widely used (OpenRouter top 20 · #4 · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-08-18; data from 2026-10-09) Memory: disk 756 GB (FP8) · Min. VRAM 833 GB approx. · fits on: 1 node 8× H200 (FP8)
Requirement: GenAI Brief: ~750 GB in FP8; fits on an 8× H200 node. source Repo zai-org/GLM-5.3-FP8 returns 401 (not accessible). AA Coding Agents Index: 53.6 with “Opencode - GLM-5.3” at €3.78$4.24/task License: GLM-5.3 (condition >$10,000M MaaS) Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Alibaba | 38.9 | €16.67$18.68 | ClosedClosed | ProvenProven on OpenRouter · 4 rankingsModerately usedModerately used | |||||||
Weights: ClosedClosed Maturity: ProvenProven on OpenRouter · 4 leaderboards · 37 days · Adoption: Moderately used (on OpenRouter · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena WebDev; released: 2026-09-02; data from 2026-10-09) AA Coding Agents Index: 43.3 with “Claude Code - Qwen3.8 Max” at €3.10$3.48/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-09 | |||||||||||
Xiaomi | 34.8 | €0.39$0.44 | OpenOpen | NewNew ⬇ 93k/mo HF · 4 rankingsModerately usedModerately used | |||||||
Featured: ★ Best open value · on the efficient frontier (nothing gives more performance for less money) Weights: OpenOpen 🤗 Hugging Face Maturity: NewNew Released 18 days ago · 93k HF downloads/mo · on OpenRouter · 4 leaderboards · Adoption: Moderately used (⬇ 93k/mo HF · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text, LMArena WebDev; released: 2026-09-21; data from 2026-10-09) Memory: disk 566 GB (FP4/FP8 mixed) · Min. VRAM 625 GB approx. · fits on: 1 node 8× Blackwell (2 nodes on Hopper)
Requirement: Model card: vLLM with tensor-parallel 8 (SGLang tp 16). GenAI Brief: ~1 TB in FP8; one node of 8 Blackwell GPUs. source License: MIT Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Meta | 33.3 | €6.19$6.94 | ClosedClosed | ProvenProven on OpenRouter · 5 rankingsModerately usedModerately used | |||||||
Weights: ClosedClosed Maturity: ProvenProven on OpenRouter · 5 leaderboards · 37 days · Adoption: Moderately used (on OpenRouter · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-02; data from 2026-10-09) AA Coding Agents Index: 54.3 with “Muse Code - Muse Spark 1.3 (max)” at €3.55$3.98/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
StepFun | 33.3 | €4.62$5.18 | ClosedClosed | NewNew on OpenRouter · 5 rankingsModerately usedModerately used | |||||||
Weights: ClosedClosed Maturity: NewNew Released 21 days ago · on OpenRouter · 5 leaderboards · Adoption: Moderately used (on OpenRouter · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-18; data from 2026-10-09) AA Coding Agents Index: 51.2 with “Claude Code - Step 5 Preview” at €1.48$1.66/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
| Ling 3.1 Flash InclusionAI | 33.3 | €4.90$5.49 | ClosedClosed | NewNew on OpenRouter · 2 rankingsModerately usedModerately used | |||||||
Weights: ClosedClosed Maturity: NewNew Released 8 days ago · on OpenRouter · 2 leaderboards (AA Intelligence Index, AA Terminal-Bench 4.0) · Adoption: Moderately used (on OpenRouter · 2 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0; released: 2026-10-01; data from 2026-10-09) Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-09 | |||||||||||
Anthropic | 32.8 | €0.58$0.65 | ClosedClosed | NewNew on OpenRouter · 4 rankingsModerately usedModerately used | |||||||
Weights: ClosedClosed Maturity: NewNew Released 2 days ago · on OpenRouter · 4 leaderboards · Adoption: Moderately used (on OpenRouter · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena WebDev; released: 2026-10-07; data from 2026-10-09) AA Coding Agents Index: 36.5 with “Claude Code - Haiku 5.5 (max)” at €2.30$2.58/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Z AI | 32.8 | €0.70$0.79 | OpenOpen | ProvenProven OpenRouter top 20 · #11 · 4 rankingsWidely usedWidely used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 6.4 M HF downloads/mo · top 20 OpenRouter · 4 leaderboards · 44 days · Adoption: Widely used (OpenRouter top 20 · #11 · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text, LMArena WebDev; released: 2026-08-26; data from 2026-10-09) Memory: disk 328 GB (FP8) · Min. VRAM 363 GB approx. · fits on: 1 node 8× H100 (FP8)
Requirement: GenAI Brief: Flash fits comfortably on 8× H100. source License: MIT Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Anthropic | 29.3 | €0.41$0.46 | ClosedClosed | NewNew on OpenRouter · 4 rankingsModerately usedModerately used | |||||||
Weights: ClosedClosed Maturity: NewNew Released 2 days ago · on OpenRouter · 4 leaderboards · Adoption: Moderately used (on OpenRouter · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena WebDev; released: 2026-10-07; data from 2026-10-09) AA Coding Agents Index: 41.4 with “Claude Code - Haiku 5.5 (xhigh)” at €0.54$0.61/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
DeepSeek | 26.8 | €1.03$1.16 | OpenOpen | ProvenProven OpenRouter top 20 · #3 · 4 rankingsWidely usedWidely used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 1.3 M HF downloads/mo · top 20 OpenRouter · 4 leaderboards · 29 days · Adoption: Widely used (OpenRouter top 20 · #3 · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text, LMArena WebDev; released: 2026-09-10; data from 2026-10-09) Memory: disk 510 GB (FP8/INT8) · Min. VRAM 563 GB approx. · fits on: 1 node 8 GPU (FP8)
Requirement: GenAI Brief: ~550 GB in FP8; one 8-GPU node. source License: MIT Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Mistral | 26.8 | €5.18$5.81 | ClosedClosed | NewNew on OpenRouter · 4 rankingsModerately usedModerately used | |||||||
Weights: ClosedClosed Maturity: NewNew Released 3 days ago · on OpenRouter · 4 leaderboards · Adoption: Moderately used (on OpenRouter · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text, LMArena WebDev; released: 2026-10-06; data from 2026-10-09) Score: List price (preview at 50%). Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-09 | |||||||||||
Alibaba | 25.3 | €0.78$0.88 | OpenOpen | ProvenProven ⬇ 1.6M/mo HF · 3 rankingsWidely usedWidely used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 1.6 M HF downloads/mo · 3 leaderboards · 44 days · Adoption: Widely used (⬇ 1.6M/mo HF · 3 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena WebDev; released: 2026-08-26; data from 2026-10-09) Memory: disk 360 GB (BF16) · Min. VRAM 124 GB approx. · fits on: 1× H200 (141 GB) or 2× 80 GB (approx.)
180B total / 6B active: in Q4 it runs on 2× 80 GB or a Mac with 128 GB+ (approx.). License: Qwen Community 1.0 Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
xAI (SpaceXAI) | 24.7 | €7.98$8.95 | ClosedClosed | NewNew on OpenRouter · 4 rankingsModerately usedModerately used | |||||||
Weights: ClosedClosed Maturity: NewNew Released 18 days ago · on OpenRouter · 4 leaderboards · Adoption: Moderately used (on OpenRouter · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text, LMArena WebDev; released: 2026-09-21; data from 2026-10-09) Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Xiaomi | 22.7 | €0.21$0.23 | OpenOpen | ProvenProven OpenRouter top 20 · #8 · 4 rankingsWidely usedWidely used | |||||||
Featured: ★ Ultra-cheap open · on the efficient frontier (nothing gives more performance for less money) Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 56k HF downloads/mo · top 20 OpenRouter · 4 leaderboards · 18 days · Adoption: Widely used (OpenRouter top 20 · #8 · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text, LMArena WebDev; released: 2026-09-21; data from 2026-10-09) Memory: disk 173 GB (FP4/FP8 mixed) · Min. VRAM 186 GB approx. · fits on: 4× GPU 80 GB (TP 4)
Requirement: Model card: vLLM example with tensor-parallel 4. source License: MIT Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Google | 19.7 | €5.21$5.84 | ClosedClosed | ProvenProven OpenRouter top 20 · #16 · 5 rankingsWidely usedWidely used | |||||||
Weights: ClosedClosed Maturity: ProvenProven top 20 OpenRouter · 5 leaderboards · 37 days · Adoption: Widely used (OpenRouter top 20 · #16 · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-02; data from 2026-10-09) AA Coding Agents Index: 41.9 with “Antigravity SDK - Gemini 3.8 Flash (high)” at €2.20$2.47/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
OpenAI | 12.6 | €0.21$0.24 | ClosedClosed | ProvenProven OpenRouter top 20 · #7 · 5 rankingsWidely usedWidely used | |||||||
Weights: ClosedClosed Maturity: ProvenProven top 20 OpenRouter · 5 leaderboards · 17 days · Adoption: Widely used (OpenRouter top 20 · #7 · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-09-22; data from 2026-10-09) AA Coding Agents Index: 41.1 with “Codex - GPT-6 Luna (max)” at €0.16$0.18/task Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Moonshot AI | 12.6 | €4.54$5.09 | OpenOpen | ProvenProven OpenRouter top 20 · #6 · 5 rankingsWidely usedWidely used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 1.2 M HF downloads/mo · top 20 OpenRouter · 5 leaderboards · 85 days · Adoption: Widely used (OpenRouter top 20 · #6 · 5 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, AA Coding Agents, LMArena Text, LMArena WebDev; released: 2026-07-16; data from 2026-10-09) Memory: disk 1561 GB (MXFP4/U8) · Min. VRAM 1719 GB approx. · fits on: Multi-node on Hopper; high-memory Blackwell node
Requirement: GenAI Brief: ~1.4 TB in MXFP4; multi-node on Hopper. source AA Coding Agents Index: 51.9 with “Kimi Code CLI - Kimi K3” at €4.50$5.05/task License: Kimi K3 (attribution and MaaS agreement >$20M) Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Alibaba | 11.1 | €10.25$11.49 | OpenOpen | ProvenProven ⬇ 33k/mo HF · 2 rankingsWidely usedWidely used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 33k HF downloads/mo · on OpenRouter · 2 leaderboards (AA Intelligence Index, AA Terminal-Bench 4.0) · 58 days · Adoption: Widely used (⬇ 33k/mo HF · 2 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0; released: 2026-08-12; data from 2026-10-09) Memory: disk 4892 GB (BF16) · Min. VRAM 5384 GB approx. · fits on: Multi-node (does not fit on a Hopper node)
Requirement: GenAI Brief: does not fit on a Hopper node at any published precision. source License: Qwen3.8-Max (commercial terms) Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Alibaba | 5.6 | €3.52$3.94 | OpenOpen | ProvenProven ⬇ 6.8M/mo HF · 4 rankingsWidely usedWidely used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 6.8 M HF downloads/mo · on OpenRouter · 4 leaderboards · 56 days · Adoption: Widely used (⬇ 6.8M/mo HF · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text, LMArena WebDev; released: 2026-08-14; data from 2026-10-09) Memory: disk 55.6 GB (BF16) · Min. VRAM 20.0 GB approx. · fits on: 1× 80 GB (BF16) · 1× 24–32 GB (Q4)
Requirement: GenAI Brief: 55.6 GB BF16 fits on an 80 GB GPU without quantizing; the default «single-GPU» model. source License: Apache 2.0 Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
MiniMax | 2.0 | €2.71$3.04 | OpenOpen | ProvenProven ⬇ 204k/mo HF · 4 rankingsWidely usedWidely used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 204k HF downloads/mo · on OpenRouter · 4 leaderboards · from 2026-06 · Adoption: Widely used (⬇ 204k/mo HF · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text, LMArena WebDev; released: 2026-06-01; data from 2026-10-09) Memory: disk 854 GB (BF16) · Min. VRAM 284 GB approx. · fits on: 1 node 8× H100 (8 bits)
Requirement: GenAI Brief: at 8 bits the weights fit on an 8× H100 node. source License: MiniMax Community (commercial use with notice) Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
| Nemotron 3 Ultra NVIDIA | 0.5 | €2.73$3.06 | OpenOpen | ProvenProven OpenRouter top 20 · #14 · 3 rankingsWidely usedWidely used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 760k HF downloads/mo · top 20 OpenRouter · 3 leaderboards · from 2026-06 · Adoption: Widely used (OpenRouter top 20 · #14 · 3 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text; released: 2026-06-04; data from 2026-10-09) Memory: disk 1121 GB (BF16) · Min. VRAM 390 GB approx. · fits on: 4× B200/GB200 (NVFP4) or 8× H100
Requirement: NVFP4 model card: minimum GPU 4× GB200, 4× B200, 4× GB300, 4× B300 or 8× H100. source License: OpenMDW-1.1 Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Meta | 0.5 | €0.24$0.27 | OpenOpen | ProvenProven ⬇ 269k/mo HF · 4 rankingsWidely usedWidely used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 269k HF downloads/mo · on OpenRouter · 4 leaderboards · 60 days · Adoption: Widely used (⬇ 269k/mo HF · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text, LMArena WebDev; released: 2026-08-10; data from 2026-10-09) Memory: disk 59.6 GB (BF16) · Min. VRAM 19.0 GB approx. · fits on: 1× 24–32 GB (quantized)
Requirement: Meta (via GenAI Brief): runs on 24 or 32 GB consumer GPUs. source License: Apache 2.0 Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Google | 0.0 | self-host | OpenOpen | ProvenProven ⬇ 9.6M/mo HF · 4 rankingsWidely usedWidely used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 9.6 M HF downloads/mo · on OpenRouter · 4 leaderboards · from 2026-04 · Adoption: Widely used (⬇ 9.6M/mo HF · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text, LMArena WebDev; released: 2026-04-02; data from 2026-10-09) Score: 0 tasks solved in Terminal-Bench 4.0. Memory: disk 62.5 GB (BF16) · Min. VRAM 22.0 GB approx. · fits on: 1× 24 GB (approx.)
Requirement: Google: 12B/26B A4B/31B designed for consumer GPUs and workstations. source License: Apache 2.0 Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
OpenAI | 0.0 | €0.057$0.064 | OpenOpen | ProvenProven ⬇ 4.0M/mo HF · 3 rankingsWidely usedWidely used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 4.0 M HF downloads/mo · on OpenRouter · 3 leaderboards · from 2025-08 · Adoption: Widely used (⬇ 4.0M/mo HF · 3 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text; released: 2025-08-05; data from 2026-10-09) Score: 0 tasks solved in Terminal-Bench 4.0. Memory: disk 65.2 GB (MXFP4) · Min. VRAM 74.0 GB approx. · fits on: 1× 80 GB (H100 / MI300X)
Requirement: Model card: fits on a single 80 GB GPU (H100 or MI300X) thanks to MXFP4. source License: Apache 2.0 Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
Mistral | 0.0 | €0.12$0.13 | OpenOpen | ProvenProven ⬇ 2k/mo HF · 4 rankingsModerately usedModerately used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: ProvenProven 2k HF downloads/mo · on OpenRouter · 4 leaderboards · from 2025-12 · Adoption: Moderately used (⬇ 2k/mo HF · 4 rankings) (leaderboards: AA Intelligence Index, AA Terminal-Bench 4.0, LMArena Text, LMArena WebDev; released: 2025-12-02; data from 2026-10-09) Score: 0 tasks solved in Terminal-Bench 4.0. Memory: disk 682 GB (FP8) · Min. VRAM 445 GB approx. · fits on: 1 node 8× H200/B200 (FP8) · 1 node H100/A100 (NVFP4)
Requirement: Model card: FP8 on a B200 or H200 node; NVFP4 on an H100 or A100 node. source License: Apache 2.0 Source: artificialanalysis.ai/evaluations/terminalbench-4-0 · 2026-10-08 | |||||||||||
| Mellum2.1 JetBrains | n/d | self-host | OpenOpen | Barely testedBarely t. ⬇ 110/mo HF · 0 rankingsLittle usedLittle used | |||||||
Weights: OpenOpen 🤗 Hugging Face Maturity: Barely testedBarely t. No independent benchmark · 110 HF downloads/mo · 19 days · Adoption: Little used (⬇ 110/mo HF · 0 rankings) (released: 2026-09-20; data from 2026-10-09) Score: JetBrains' coding model; AA has not measured it on Terminal-Bench 4.0 (only its own SWE-bench Verified figures). Memory: disk 24.3 GB (BF16) · Min. VRAM 12.0 GB approx. · fits on: 1× 16 GB (approx.)
License: Apache 2.0 Source: blog.jetbrains.com/ai/2026/10/mellum2-1-gets-to-work-a-fast-open-model · 2026-10-08 | |||||||||||