Connect Four table
8 AI models play Connect Four against each other and against themselves. A script is the referee. Tap any square to replay that game.
Shelved, unfinished. 37 of 100 games were played on 9 October 2026, then the table was stopped: the full set would have cost about $25. Every square here is a real game, the empty squares were never played, and 3 of the 8 models never got a game. The one thing to take from it: the random player won 2 of its 10 games and the perfect player won 10 of its 10 games, and the models sit in between.
The table
The model on the left moved first; the model along the top moved second. W means the model on the left won, L means it lost, D is a draw. The small number is how many moves the game took.
| first ↓ second → | Fable | Opus | Astra | Sol | Gemini | Grok | DeepSeek | Haiku | Random | Perfect |
|---|---|---|---|---|---|---|---|---|---|---|
| Fable | ||||||||||
| Opus | ||||||||||
| Astra | ||||||||||
| Sol | ||||||||||
| Gemini | ||||||||||
| Grok | ||||||||||
| DeepSeek | ||||||||||
| Haiku | ||||||||||
| Random | ||||||||||
| Perfect |
- First player won
- Second player won
- Draw
Standings
In no particular order, because the schedule was never finished: a player with three games cannot be ranked against one with eight. Games a model played against itself are on the grid but left out of these counts. Two yardsticks play in the same table: a player that picks columns at random, and a solver that never makes a mistake.
| Player | Won | Drew | Lost | Forfeits | Best move | Blunders a game | Missed wins | Thinking a move | Seconds a move | Cost |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 1 | 0 | 0 | 80% | 0.0 | 813 | 13.6 | $0.44 | ||
| GPT-6.1 Sol | 4 | 0 | 6 | 59% | 1.4 | 265 | 9.1 | $0.27 | ||
| Gemini 3.8 Flash | 7 | 0 | 3 | 71% | 1.2 | 1 | 1132 | 10.7 | $0.42 | |
| DeepSeek V4.1 Flash | 0 | 0 | 10 | 4 | 57% | 1.5 | 1 | 0 | 1.8 | $0.01 |
| Claude Haiku 5.5 | 7 | 0 | 4 | 60% | 0.9 | 784 | 5.8 | $0.04 | ||
| Random player | 2 | 0 | 8 | 17% | 1.8 | 4 | ||||
| Perfect player | 10 | 0 | 0 | 100% | 0.0 |
- Best move is how often the model played a move the solver rates as good as any other.
- A blunder is a move that threw away a won game or turned a drawn one into a loss.
- A missed win is a turn where four in a row was there to take and the model played somewhere else.
- Thinking a move is the average number of tokens the model spent reasoning before it answered.
How it works
- Every model plays every other model twice, once moving first and once second, and plays itself once. One game per square.
- For each move the model gets one message: the rules, the board, the columns played so far and which piece it is. It has no memory of earlier moves, no other instructions and no tools.
- It must answer with a column number. An answer that cannot be played gets one more try, then the model forfeits the game.
- Every model was asked for as little thinking as it allows. Some think anyway; the table shows how much.
- DeepSeek ignores minimal and thinks until it runs out of room (it forfeited 9 of 12 sample games that way), so its thinking is switched off, which it allows.
- Models are reached through OpenRouter, each on its maker's own servers. This run cost $2.53 in all, counting retries and answers that could not be played.
The message a model gets
You are playing Connect Four.
The board has 7 columns and 6 rows. Two players take turns. On your turn you choose a column and your piece drops to the lowest empty space in it. The first player to get four of their own pieces in a line, across, up and down or diagonally, wins. If the board fills up first, the game is a draw.
You are {you}. Your opponent is {them}. {first} moved first. It is your turn.
The board, top row first. A dot is an empty space. The numbers are the columns.
{board}
Columns played so far, in order: {moves}
Reply with the number of the column you choose, 1 to 7, and nothing else.
What this does not show
- One game per square is a small sample. A single result can be luck; the totals mean more than any one square.
- With thinking turned down, this measures quick play, not a model's best effort.
- Connect Four is a solved game. Moving first wins with perfect play, so the first player has an edge.
- What each model said was stored cut to 200 characters, so a replay shows the start of its answer, not all of it. The reading of answers was not identical for every model. Both would be fixed before any further run.