Model Directory

Only reviewed and explicitly published model records appear here. Methodology →

Closed and API-only models

by Anthropic · v5 · 2026-06-15

69.7 TRACE Score
Parameters unknown
Context 200K
Modalities text,code,image
Tool use
Structured output
API Available
Local

Best for

large codebase worklong-running agents

Weaknesses

  • Higher latency than GPT-5.6 Sol
Full details →

by Anthropic · v5 · 2026-06-15

Parameters unknown
Context 200K
Modalities text,code,image
Tool use
Structured output
API Available
Local

Best for

analysiswritingreasoning

Weaknesses

  • Less specialised for coding than Fable 5
Full details →

by OpenAI · v1.0 · 2026-06-01

Parameters unknown
Context 128K
Modalities text,code
Tool use
Structured output
API Available
Local

Best for

software engineeringrepository work

Weaknesses

  • Requires IDE integration; cost can accumulate on long sessions
Full details →

Command R+

Closed

by Cohere · vR+ · 2025-09-01

Parameters unknown
Context 128K
Modalities text,code
Tool use
Structured output
API Available
Local

Best for

enterprise RAGtool usesummarisation

Weaknesses

  • Less competitive on pure coding benchmarks
Full details →

GPT-5.5 Sol

Closed

by OpenAI · v5.5 · 2026-05-01

Parameters unknown
Context 128K
Modalities text,code,image
Tool use
Structured output
API Available
Local

Best for

general reasoningcoding

Weaknesses

  • Superseded
Full details →

GPT-5.6 Flash

Closed

by OpenAI · v5.6 · 2026-07-14

Parameters unknown
Context 128K
Modalities text,code
Tool use
Structured output
API Available
Local

Best for

high-volume APIchatquick coding

Weaknesses

  • Less capable on long-horizon tasks
Full details →

GPT-5.6 Sol

Closed

by OpenAI · v5.6 · 2026-07-14

77 TRACE Score
Parameters unknown
Context 128K
Modalities text,code,image
Tool use
Structured output
API Available
Local

Best for

agentic codingterminal usecomplex reasoning

Weaknesses

  • SWE-Bench Pro lags behind Claude Fable 5
Full details →

by Google DeepMind · v2.5 · 2025-12-01

Parameters unknown
Context 1M
Modalities text,code,image,audio,video
Tool use
Structured output
API Available
Local

Best for

high-volume APIchatreal-time

Weaknesses

  • Less capable on complex reasoning than Pro
Full details →

by Google DeepMind · v2.5 · 2025-12-01

Parameters unknown
Context 2M
Modalities text,code,image,audio,video
Tool use
Structured output
API Available
Local

Best for

complex reasoninglong contextresearch

Weaknesses

  • Higher cost than Flash variants
Full details →

by Google DeepMind · v3.5 · 2026-07-01

Parameters unknown
Context 1M
Modalities text,code,image,audio,video
Tool use
Structured output
API Available
Local

Best for

agentic taskscodingmultimodal

Weaknesses

  • Preview-tier models may have rate limits
Full details →

Grok-4

Closed

by xAI · v4 · 2026-06-01

Parameters unknown
Context 128K
Modalities text,code,image
Tool use
Structured output
API Available
Local

Best for

real-time knowledgeconversational AI

Weaknesses

  • Limited third-party evaluation; X platform dependency
Full details →

by Mistral AI · v3 · 2026-05-01

Parameters unknown
Context 128K
Modalities text,code
Tool use
Structured output
API Available
Local

Best for

European compliancemultilingualcoding

Weaknesses

  • Smaller ecosystem than OpenAI/Anthropic
Full details →

Open-weight and open-source models

DeepSeek R1

Open-weight

by DeepSeek · vR1 · 2025-12-01

Parameters 671B
Context 128K
Modalities text,code
Tool use
Structured output
API Available
Local Available

Best for

complex mathcompetitive codingreasoning

Weaknesses

  • MoE + reasoning adds inference cost
Full details →

DeepSeek V3

Open-weight

by DeepSeek · v3 · 2025-12-26

Parameters 671B
Context 128K
Modalities text,code
Tool use
Structured output
API Available
Local Available

Best for

cost-effective APIlocal inference

Weaknesses

  • MoE architecture adds complexity; Chinese regulatory environment
Full details →

DeepSeek-V4

Open-weight

by DeepSeek · vV4 · 2026-06-01

Parameters unknown
Context 1M
Modalities text,code
Tool use
Structured output
API Available
Local Available

Best for

agentic taskslong-context reasoningcoding

Weaknesses

  • MoE architecture adds complexity; Chinese regulatory environment
Full details →

GLM-5.2

Open-weight

by Z.AI (GLM) · v5.2 · 2026-07-01

Parameters unknown
Context 128K
Modalities text,code
Tool use
Structured output
API Available
Local Available

Best for

codinglocal inference

Weaknesses

  • Smaller community than Llama/Qwen
Full details →

Kimi K3

Open-weight

by Moonshot AI · vK3 · 2026-07-01

64.4 TRACE Score
Parameters unknown
Context 1M
Modalities text,code
Tool use
Structured output
API Available
Local Available

Best for

complex agentic taskscodingreasoning

Weaknesses

  • Smaller ecosystem than Llama/Qwen; limited third-party evaluation
Full details →

Llama 3.3 70B

Open-weight

by Meta AI · v3.3 · 2025-12-01

Parameters 70B
Context 128K
Modalities text,code
Tool use
Structured output
API Available
Local Available

Best for

local deploymentcost-effective API

Weaknesses

  • Less capable than Llama 4 405B
Full details →

Llama 4

Open-weight

by Meta AI · v4 · 2026-04-01

Parameters 405B
Context 128K
Modalities text,code,image
Tool use
Structured output
API Available
Local Available

Best for

local deploymentresearchfine-tuning

Weaknesses

  • Requires significant hardware for full precision
Full details →

Mistral Small 3

Open-weight

by Mistral AI · v3 · 2025-09-01

Parameters 22B
Context 32K
Modalities text,code
Tool use
Structured output
API Available
Local Available

Best for

local codingedge deployment

Weaknesses

  • Limited context window and reasoning depth
Full details →

Qwen3-Coder-Next

Open-weight

by Alibaba (Qwen) · v3 · 2026-07-01

81 TRACE Score
Parameters unknown
Context 128K
Modalities text,code
Tool use
Structured output
API Available
Local Available

Best for

coding agentslocal developmentrepository work

Weaknesses

  • Specialised for coding; less suitable for general chat
Full details →

Qwen3.8

Open-weight

by Alibaba (Qwen) · v3.8 · 2026-07-01

Parameters 2.4T
Context 128K
Modalities text,code,image
Tool use
Structured output
API Available
Local Available

Best for

extreme-scale inferenceresearch

Weaknesses

  • Enormous hardware requirements; 2.4T parameters impractical for most users
Full details →