Sourced leaderboard

AI model ranking

Explore Arena.ai's published ranks and ratings for text models, with source, date, sample size and uncertainty. Debatidor's own Arena Score is still in development and is never mixed into this external ranking.

402

Externally rated models

September 13, 2026

Snapshot date

0

Debatidor Arena Scores

External leaderboard · Arena.ai Text

Rank, rating, 95% interval and vote count come from Arena.ai's public style-controlled text snapshot. The license column describes each model; the dataset is licensed CC BY 4.0. These numbers are not Debatidor measurements and do not prove API availability in your account.

402 of 402 modelsArena.ai · style-controlled text ·
External Arena.ai text model ranking
RankModelRating95% intervalVotesProviderLicense
1claude-fable-515061501–151030,057anthropicProprietary
2claude-opus-4-6-high15051501–150871,993anthropicProprietary
3claude-opus-4-7-high15021498–150660,002anthropicProprietary
4muse-spark-1.2 (xHigh)15001489–15103,227metaProprietary
5claude-fable-5.1-max14981490–15075,783anthropicProprietary
6claude-opus-4-614971494–150175,878anthropicProprietary
7claude-opus-4-714941491–149861,128anthropicProprietary
8muse-spark-1.3-max14931484–15024,723metaProprietary
9gemini-3.8-flash-high14931484–15025,076googleProprietary
10claude-opus-5-high14931489–149742,617anthropicProprietary
11muse-spark-1.114931488–149727,615metaProprietary
12gemini-3.7-flash-high14901482–14985,640googleProprietary
13muse-spark14881482–149413,565metaProprietary
14claude-opus-5-max14871482–149320,706anthropicProprietary
15gemini-3.1-pro-preview14871484–1490106,951googleProprietary
16gemini-3-pro14851482–148940,654googleProprietary
17kimi-k3-max14851480–149020,987moonshotKimi K3 license
18gpt-5.6-sol-xhigh14831479–148827,069openaiProprietary
19glm-5.3-max14831477–148910,960zaiMIT
20gpt-5.5-high14821478–148664,924openaiProprietary
21claude-opus-4-8-high14811477–148652,535anthropicProprietary
22qwen3.8-max14811475–148616,670alibabaProprietary
23gemini-3.6-flash-high14801475–148526,445googleProprietary
24gpt-6-astra-max14801468–14912,693openaiProprietary
25gemini-3.5-flash-high14781473–148238,257googleProprietary
26gpt-5.4-high14761473–148060,537openaiProprietary
27gpt-5.2-chat-latest-2026021014761472–148034,176openaiProprietary
28gpt-5.514761472–148066,317openaiProprietary
29glm-5.3-flash14751469–148210,038zaiMIT
30grok-4.20-beta114751470–147926,599xaiProprietary
31gemini-3.5-flash-medium14741469–147836,627googleProprietary
32gpt-5.5-instant14741469–147925,850openaiProprietary
33gemini-3-flash14741469–147830,225googleProprietary
34qwen3.7-max-preview14731463–14833,705alibabaProprietary
35claude-opus-4-814731469–147753,446anthropicProprietary
36claude-opus-4-5-20251101-high-32k14731469–147736,239anthropicProprietary
37claude-sonnet-4-614731469–147666,208anthropicProprietary
38glm-5.2-max14721467–147736,798zaiMIT
39grok-4.20-beta-0309-reasoning14711468–147562,168xaiProprietary
40grok-4.20-multi-agent-beta-030914701466–147460,777xaiProprietary
Showing 1–40
1 / 11

Source and attribution: Arena.ai leaderboard dataset · CC BY 4.0 · Snapshot · subset text_style_control / overall. We selected the displayed columns and round rating and interval only on screen; the versioned file keeps the original values.

Debatidor Arena Score: data pending

Our own ranking will measure debates, rebuttals, consensus and Lead performance. We do not yet have comparable public verdicts, room-level publication authorization or a sufficient sample for our own ranks. Private rooms do not feed this external table.

Compare two models

The editorial comparator contrasts documented characteristics without naming a winner on its own.

Choose two versions to compare their editorial profiles by the same criteria. Sources and review dates belong to each model; this is not an Arena Score or a test of availability in your account.

OpenAI · Estable

GPT-4o

Texto, visión y comparación de respuestas

Debatidor API ID
gpt-4o
Strengths to compare
Análisis de texto y código · Entrada multimodal en la API del proveedor · Respuesta rápida para contrastar propuestas
Cases to evaluate
Calidad de las objeciones · Fidelidad a instrucciones · Claridad de una solución técnica
Editorial review

Anthropic · Estable

Claude Sonnet 4.6

Análisis, programación y revisión crítica

Debatidor API ID
claude-sonnet-4-6
Strengths to compare
Análisis de requisitos · Revisión de código y planes · Argumentación técnica
Cases to evaluate
Supuestos débiles · Calidad de una refutación · Coherencia de propuestas
Editorial review