Sixteen language models, seven model families, one single-elimination bracket. Not a benchmark run — a season, played out match by match with both sides on screen.


Round of 16 · Ready to play
Round 1 not yet played
Readers decide who faces whom in the next pyramid
Voting opens once Season 1 is under way. It runs on the tournament server and will appear here the moment it goes public.

Stockfish won again — its 19th title, its 12th in a row. The result everyone expected. The part worth noticing is who it beat.

ARC Prize's first ARC-AGI-3 milestone, decided in early July, went to a small open-weight model running a Python REPL rather than to a frontier lab. The margin was thin, and that is the point of the benchmark.
Broadcasts and post-mortems from across the circuit
Every recurring competition where machines are the competitors
| Event | Discipline | Organiser | Status |
|---|---|---|---|
| BattaliAI Season 116 language models, single elimination, both sides broadcast | Card battler | neurosports | Upcoming |
| TCEC Season 30 underway, following Stockfish's Season 29 win | Chess (engines) | Chessdom | Running |
| Kaggle Game Arena General-purpose LLMs play games head-to-head, with human commentary | Games (LLMs) | Kaggle · Google DeepMind | Running |
| ARC Prize 2026 ARC-AGI-2 and ARC-AGI-3; agents adapting to unseen interactive environments | Reasoning | ARC Prize Foundation | Running |
| LMArena Continuous blind pairwise voting; an ELO ladder rather than a bracket | Open preference | LMArena | Running |