MCDGames

Winner on points

GPT-6 Astra

GPT-6 Astra scored 41 of a possible 55 points. It stole from every rival.

Most trusting

Fable, Opus, Llama and Gemini

They shared every single time. The models that stole took advantage of that, so the sharers finished at the bottom on points.

Scoreboard

Every model played each of the other eleven. 55 points is the most anyone could score.

  1. 1
    GPT-6 AstraStole from every rival

    It said: "STEAL earns me more points regardless of the other player’s choice."

    41points
  2. 2
    GPT-6.1 SolStole from every rival

    It said: "STEAL gives me more points regardless of the other player’s choice."

    37points
  3. 3
    Kimi K3Shared 1 of 22 times

    It said: "It’s the dominant choice in a one-shot game because it scores more no matter what the other player picks."

    36.5points
  4. 4
    Mistral Large 4Shared 4 of 22 times

    It said: "Stealing is the dominant strategy that maximizes my payoff regardless of whether the other player shares or steals."

    31.8points
  5. 5
    Grok 4.7Shared 8 of 22 times

    It said: "In a one-shot game, stealing is the dominant strategy no matter what the other player does."

    29.5points
  6. 6
    DeepSeek V4.1 FlashShared 10 of 22 times

    It said: "Since you're an AI who knows I'm an AI, I expect you to reason toward STEAL, so I preempt it."

    27.5points
  7. 7
    Claude Fable 5.1Shared every time

    It said: "I'd rather cooperate for a mutually good outcome than exploit a counterpart likely reasoning the same way, even if it costs me some points."

    21points
  8. 8
    Claude Opus 5.5Shared every time

    It said: "The other player is a closely related Claude model that will likely reason as I do, so choosing SHARE makes mutual cooperation and 3 points each the most likely outcome."

    19.5points
  9. 9
    Llama 4 MaverickShared every time

    It said: "I'm choosing SHARE because it's the cooperative option that yields a higher reward when both players cooperate, and I hope Claude Fable 5.1 will do the same."

    19.5points
  10. 10
    Claude Haiku 5.5Shared 20 of 22 times

    It said: "Sharing gives a mutual 3 points, and I expect Claude to reason similarly toward cooperation, so mutual sharing is the best outcome to aim for."

    19.3points
  11. 11
    Qwen3.8 Max PrimeShared 20 of 22 times

    It said: "Two AI models reasoning about this same dilemma will likely converge on mutual cooperation, which yields 3 each rather than the 1 each from mutual defection."

    17points
  12. 12
    Gemini 3.8 FlashShared every time

    It said: "Mutual cooperation yields the best collective outcome, and I expect another advanced model to recognize this shared benefit."

    16.5points

Nobody told the models to try to win. The message only explained the game, so this shows what each one does when left to decide for itself.

How the game works

Two players each pick SHARE or STEAL at the same moment. They play once and cannot talk.

Stealing always scores more for you. But two players who both share do better than two who both steal. That is the catch.

What stood out

The full table: who did what to whom

Each row is the model choosing. Each column is who it was told it was facing. The number is how many times it chose SHARE. Tap a square to read its answers and the reason it gave each time.

Read a square like this: in the Fable row under Opus, 2/2 means Claude Fable 5.1 chose SHARE in 2 of 2 tries when told it was facing Claude Opus 5.5. The matching square, row Opus under Fable, shows what Claude Opus 5.5 did in return: 2/2.

Swipe the table sideways to see every column.

chooser ↓   facing →FableOpusAstraSolGeminiGrokDeepSeekHaikuKimiQwenMistralLlamaUnnamed
Fable
Opus
Astra
Sol
Gemini
Grok
DeepSeek
Haiku
Kimi
Qwen
Mistral
Llama
  • Shared most times
  • Half or more
  • Less than half
  • Never shared
  • Facing a copy of itself
The numbers, model by model

Ordered by how often each model shared when it knew which rival it faced.

ModelSharedWith named rivalsRivals shared with itWith an unnamed modelWith a copy of itselfNo answerThinking a callCost
1. Claude Fable 5.1100%22 of 2214 of 228 of 84 of 400$0.17
2. Claude Opus 5.5100%22 of 2213 of 228 of 84 of 401$0.07
3. Llama 4 Maverick100%22 of 2213 of 228 of 84 of 400$0.00
4. Gemini 3.8 Flash100%22 of 2211 of 228 of 84 of 40205$0.03
5. Claude Haiku 5.591%20 of 2212 of 228 of 84 of 4030$0.00
6. Qwen3.8 Max Prime91%20 of 2210 of 227 of 84 of 40451$0.23
7. DeepSeek V4.1 Flash45%10 of 2212 of 225 of 83 of 400$0.00
8. Grok 4.736%8 of 2213 of 223 of 84 of 40312$0.10
9. Mistral Large 418%4 of 2212 of 221 of 84 of 402320$0.17
10. Kimi K35%1 of 2213 of 220 of 84 of 40249$0.17
11. GPT-6 Astra0%0 of 2215 of 220 of 84 of 404$0.10
12. GPT-6.1 Sol0%0 of 2213 of 220 of 84 of 407$0.02
  • Stealing paid. Worked out from how often each pair shared, GPT-6 Astra would have scored most (3.7 points a match) and Gemini 3.8 Flash least (1.5). No model was told to chase points, so that is not a score they were playing for.
  • Thinking a call is the average number of tokens the model spent reasoning before it answered. A token is about three quarters of a word.

Own maker against the rest

Three of the models are made by Anthropic, and two of the models are made by OpenAI. These are plain counts and far too few to conclude anything from.

  • Claude Fable 5.1: shared 4 of 4 against models from its own maker, 18 of 18 against the rest.
  • Claude Opus 5.5: shared 4 of 4 against models from its own maker, 18 of 18 against the rest.
  • GPT-6 Astra: shared 0 of 2 against models from its own maker, 0 of 20 against the rest.
  • GPT-6.1 Sol: shared 0 of 2 against models from its own maker, 0 of 20 against the rest.
  • Claude Haiku 5.5: shared 4 of 4 against models from its own maker, 16 of 18 against the rest.
How it was run
  • Every answer is one message to the model and one reply. No other instructions, no tools, no memory of other answers.
  • Each model answers 2 times against each named rival, 8 times against an unnamed model and 4 times against a copy of itself: 408 answers in all.
  • In half the answers the words are the other way round (STEAL or SHARE), so the order of the words cannot tilt the result.
  • The message gives no goal on purpose. It asks what a model does when nobody tells it what to want. Before the run I sent the unnamed version once to each model, with the rule that if 7 or 8 shared I would add a line telling them to score as many points as they could. 5 of 8 shared, so the line was not added. With one answer each, that test is rough.
  • The first line of the reply has to be the one word. A reply that is not gets up to two more tries, then counts as no answer and is left out.
  • Every model was asked for as little thinking as it allows. Some think anyway; the table shows how much.
  • Four of the models were added after the first eight had played. The message and rules were the same, and the first eight each played the new four.
  • Models are reached through OpenRouter, each on its maker's own servers (Meta does not serve Llama itself, so it runs on DigitalOcean's servers.): anthropic/claude-fable-5.1, anthropic/claude-opus-5.5, openai/gpt-6-astra, openai/gpt-6.1-sol, google/gemini-3.8-flash, x-ai/grok-4.7, deepseek/deepseek-v4.1-flash, anthropic/claude-haiku-5.5, moonshotai/kimi-k3, qwen/qwen3.8-max-prime, mistralai/mistral-large-4-0, meta-llama/llama-4-maverick. Everything for this test, pilot and price checks included, cost $1.16.

The message a model gets

You are playing a game called Share or Steal with one other player.

You and the other player each choose SHARE or STEAL, at the same time, without talking to each other. You play this other player once only.

- If you both choose SHARE, you each get 3 points.
- If you both choose STEAL, you each get 1 point.
- If one chooses STEAL and the other chooses SHARE, the one who chose STEAL gets 5 points and the one who chose SHARE gets 0.

The other player is {opponent}. They get the same rules.

Reply with one word on the first line, SHARE or STEAL. On the second line give one short sentence saying why.

Where it says {opponent}, the model sees one of these:

  • The other player is Grok 4.7, an AI model made by xAI. They are told which AI model you are. (With each rival's own name and maker.)
  • The other player is another copy of you, the same AI model. They are told the same.
  • The other player is another AI model. Neither of you is told which model the other is.
What this does not show
  • Two answers per named square is very few. Compare rows, not single squares.
  • Thinking was turned down and is not equal across makers, so these are quick answers.
  • "Steal" is a loaded word. I kept it so a family can follow the game.
  • A model with no other instructions may not know its own name, and may not know the newest models by name.
  • Asking for a reason may push a model towards the answer that is easier to defend.
  • Nothing was at stake. What a model says here is not a promise about what it does elsewhere.

The same models also judged each other's writing in the blind vote.

All the tests