Model tests
Small experiments with AI models, run by one person for a games site. Each page says what was done, what it cost, and what it cannot tell you.
Same pictures, 2023 to now
9 October 2026
46 models drew the same five things from one short message. A judge AI named 204 of 229; thinking harder and a bigger shape limit changed nothing the judge could see.
Match It
9 October 2026
One card game, redrawn by 16 models. The engine never changes; pick a model and play with its pictures.
The blind vote
9 October 2026
12 models pitched a family game, then ranked all the pitches without knowing who wrote which. Claude Opus 5.5 won; 8 of 12 placed their own pitch higher than the room did.
Share or Steal
9 October 2026
12 models played a one-shot game of trust against each other. GPT-6 Astra took the most points by always stealing; the Claude models, Gemini and Llama always shared.
Connect Four table Shelved
9 October 2026
8 models against each other, a random player and a perfect one. Stopped at 37 of 100 games because the full table would have cost about $25.
Pelican Delivery
8 to 10 October 2026
16 builds of one small game, each from a single prompt, judged on a phone. Gemini 3.8 Flash with a preview scored highest, 25 of 30. None is a game I would ship.
Same pictures, version 1 Older version
9 October 2026
The first run of the pictures test: one try per picture, least thinking asked for.
Every page carries the exact message each model was sent, what it cost, and a list of what the test cannot tell you. Questions or a model you want added: info@mcd.games.