Sourced leaderboard
AI model ranking
Explore Arena.ai's published ranks and ratings for text models, with source, date, sample size and uncertainty. Debatidor's own Arena Score is still in development and is never mixed into this external ranking.
402
Externally rated models
September 13, 2026
Snapshot date
0
Debatidor Arena Scores
External leaderboard · Arena.ai Text
Rank, rating, 95% interval and vote count come from Arena.ai's public style-controlled text snapshot. The license column describes each model; the dataset is licensed CC BY 4.0. These numbers are not Debatidor measurements and do not prove API availability in your account.
| Rank | Model | Rating | 95% interval | Votes | Provider | License |
|---|---|---|---|---|---|---|
| 1 | claude-fable-5 | 1506 | 1501–1510 | 30,057 | anthropic | Proprietary |
| 2 | claude-opus-4-6-high | 1505 | 1501–1508 | 71,993 | anthropic | Proprietary |
| 3 | claude-opus-4-7-high | 1502 | 1498–1506 | 60,002 | anthropic | Proprietary |
| 4 | muse-spark-1.2 (xHigh) | 1500 | 1489–1510 | 3,227 | meta | Proprietary |
| 5 | claude-fable-5.1-max | 1498 | 1490–1507 | 5,783 | anthropic | Proprietary |
| 6 | claude-opus-4-6 | 1497 | 1494–1501 | 75,878 | anthropic | Proprietary |
| 7 | claude-opus-4-7 | 1494 | 1491–1498 | 61,128 | anthropic | Proprietary |
| 8 | muse-spark-1.3-max | 1493 | 1484–1502 | 4,723 | meta | Proprietary |
| 9 | gemini-3.8-flash-high | 1493 | 1484–1502 | 5,076 | Proprietary | |
| 10 | claude-opus-5-high | 1493 | 1489–1497 | 42,617 | anthropic | Proprietary |
| 11 | muse-spark-1.1 | 1493 | 1488–1497 | 27,615 | meta | Proprietary |
| 12 | gemini-3.7-flash-high | 1490 | 1482–1498 | 5,640 | Proprietary | |
| 13 | muse-spark | 1488 | 1482–1494 | 13,565 | meta | Proprietary |
| 14 | claude-opus-5-max | 1487 | 1482–1493 | 20,706 | anthropic | Proprietary |
| 15 | gemini-3.1-pro-preview | 1487 | 1484–1490 | 106,951 | Proprietary | |
| 16 | gemini-3-pro | 1485 | 1482–1489 | 40,654 | Proprietary | |
| 17 | kimi-k3-max | 1485 | 1480–1490 | 20,987 | moonshot | Kimi K3 license |
| 18 | gpt-5.6-sol-xhigh | 1483 | 1479–1488 | 27,069 | openai | Proprietary |
| 19 | glm-5.3-max | 1483 | 1477–1489 | 10,960 | zai | MIT |
| 20 | gpt-5.5-high | 1482 | 1478–1486 | 64,924 | openai | Proprietary |
| 21 | claude-opus-4-8-high | 1481 | 1477–1486 | 52,535 | anthropic | Proprietary |
| 22 | qwen3.8-max | 1481 | 1475–1486 | 16,670 | alibaba | Proprietary |
| 23 | gemini-3.6-flash-high | 1480 | 1475–1485 | 26,445 | Proprietary | |
| 24 | gpt-6-astra-max | 1480 | 1468–1491 | 2,693 | openai | Proprietary |
| 25 | gemini-3.5-flash-high | 1478 | 1473–1482 | 38,257 | Proprietary | |
| 26 | gpt-5.4-high | 1476 | 1473–1480 | 60,537 | openai | Proprietary |
| 27 | gpt-5.2-chat-latest-20260210 | 1476 | 1472–1480 | 34,176 | openai | Proprietary |
| 28 | gpt-5.5 | 1476 | 1472–1480 | 66,317 | openai | Proprietary |
| 29 | glm-5.3-flash | 1475 | 1469–1482 | 10,038 | zai | MIT |
| 30 | grok-4.20-beta1 | 1475 | 1470–1479 | 26,599 | xai | Proprietary |
| 31 | gemini-3.5-flash-medium | 1474 | 1469–1478 | 36,627 | Proprietary | |
| 32 | gpt-5.5-instant | 1474 | 1469–1479 | 25,850 | openai | Proprietary |
| 33 | gemini-3-flash | 1474 | 1469–1478 | 30,225 | Proprietary | |
| 34 | qwen3.7-max-preview | 1473 | 1463–1483 | 3,705 | alibaba | Proprietary |
| 35 | claude-opus-4-8 | 1473 | 1469–1477 | 53,446 | anthropic | Proprietary |
| 36 | claude-opus-4-5-20251101-high-32k | 1473 | 1469–1477 | 36,239 | anthropic | Proprietary |
| 37 | claude-sonnet-4-6 | 1473 | 1469–1476 | 66,208 | anthropic | Proprietary |
| 38 | glm-5.2-max | 1472 | 1467–1477 | 36,798 | zai | MIT |
| 39 | grok-4.20-beta-0309-reasoning | 1471 | 1468–1475 | 62,168 | xai | Proprietary |
| 40 | grok-4.20-multi-agent-beta-0309 | 1470 | 1466–1474 | 60,777 | xai | Proprietary |
Source and attribution: Arena.ai leaderboard dataset · CC BY 4.0 · Snapshot · subset text_style_control / overall. We selected the displayed columns and round rating and interval only on screen; the versioned file keeps the original values.
Debatidor Arena Score: data pending
Our own ranking will measure debates, rebuttals, consensus and Lead performance. We do not yet have comparable public verdicts, room-level publication authorization or a sufficient sample for our own ranks. Private rooms do not feed this external table.
Compare two models
The editorial comparator contrasts documented characteristics without naming a winner on its own.
Choose two versions to compare their editorial profiles by the same criteria. Sources and review dates belong to each model; this is not an Arena Score or a test of availability in your account.
OpenAI · Estable
GPT-4o
Texto, visión y comparación de respuestas
- Debatidor API ID
- gpt-4o
- Strengths to compare
- Análisis de texto y código · Entrada multimodal en la API del proveedor · Respuesta rápida para contrastar propuestas
- Cases to evaluate
- Calidad de las objeciones · Fidelidad a instrucciones · Claridad de una solución técnica
- Editorial review
Anthropic · Estable
Claude Sonnet 4.6
Análisis, programación y revisión crítica
- Debatidor API ID
- claude-sonnet-4-6
- Strengths to compare
- Análisis de requisitos · Revisión de código y planes · Argumentación técnica
- Cases to evaluate
- Supuestos débiles · Calidad de una refutación · Coherencia de propuestas
- Editorial review