Pelican Delivery
I give AI models the same prompt and one try to build the same small game. No help, no second attempt, no fixing. Then I play every build on my phone and score it.
Test 1: Pelican Delivery
A pelican cycles a basket of fish over the hills to market before sunset. One control: hold to pedal, let go to coast. 8 models built it. Tap Play to try any of them; they are exactly as the models returned them. Built for a phone held upright.
-
1GPT-6.1 Sol
OpenAI
24 out of 30
- Checks
- 14/15
- Look
- 4/5
- Feel
- 3/5
- Fun
- 3/5
$0.18 to build. 18,382 tokens in 6 min.
Play -
2GPT-6 Astra
OpenAI
23 out of 30
- Checks
- 15/15
- Look
- 4/5
- Feel
- 2/5
- Fun
- 2/5
$1.04 to build. 20,792 tokens in 6 min.
Play -
3Gemini 3.8 Flash
Google
21 out of 30
- Checks
- 15/15
- Look
- 1/5
- Feel
- 2/5
- Fun
- 3/5
$0.05 to build. 12,861 tokens in 87 sec.
Play -
4Claude Opus 5.5
Anthropic
20 out of 30
- Checks
- 14/15
- Look
- 3/5
- Feel
- 2/5
- Fun
- 1/5
$1.77 to build. 88,326 tokens in 14 min.
Play -
5Grok 4.7
xAI
19 out of 30
- Checks
- 15/15
- Look
- 1/5
- Feel
- 2/5
- Fun
- 1/5
$0.71 to build. 118,343 tokens in 28 min.
Play -
6Claude Haiku 5.5
Anthropic
18 out of 30
- Checks
- 14/15
- Look
- 2/5
- Feel
- 1/5
- Fun
- 1/5
$0.04 to build. 76,021 tokens in 6 min.
Play -
7DeepSeek V4.1 Flash
DeepSeek
17 out of 30
- Checks
- 14/15
- Look
- 1/5
- Feel
- 1/5
- Fun
- 1/5
$0.04 to build. 66,107 tokens in 4 min.
Play -
8Claude Fable 5.1
Anthropic
16 out of 30
- Checks
- 13/15
- Look
- 1/5
- Feel
- 1/5
- Fun
- 1/5
$4.93 to build. 98,483 tokens in 20 min.
Play
Prompt v1. Built and judged 8 to 10 October 2026. One run per model.
What stood out
My honest take: None of these is a game I would put on mcd.games. Some were better than others, but the best of them is only okay.
- GPT-6.1 Sol came first with 24 out of 30. It cost $0.18 to build.
- Price did not predict the result. The most expensive build, Claude Fable 5.1 at $4.93, came eighth. The cheapest, Claude Haiku 5.5 at $0.04, came sixth.
- Following the rules was the easy part. 7 of 8 builds passed at least 14 of the 15 checks. Only 3 got more than a single star for fun.
- Most of the effort is invisible. Across all 8 builds, 73% of the tokens were spent thinking before any code was written.
How the test works
- One message. Every model gets the prompt below and nothing else: no system prompt, no tools, no follow-up. It cannot run its own code or see what it made.
- One reply. Whatever comes back is the build. I never edit it.
- The same route for all. Each model is called through OpenRouter, served by the company that makes it, at its default effort.
- Judged by playing. A script checks four things that need no opinion. I play three rounds of each build on a phone and answer eleven yes or no questions, then give one to five stars each for look, feel and fun. That makes 30 points.
- Judged blind. The builds are shown to me as Build A, Build B and so on.
The prompt
This is the whole message, word for word.
Build a small browser game called Pelican Delivery. The game: a pelican rides a bicycle from left to right across rolling hills, carrying a basket of 10 fish to the fish market at the end of the road. The market closes at sunset, about 60 seconds after the round starts. The rules: - One control only. Touch and hold anywhere to pedal. Let go to coast. Coasting gently slows the bike, so letting go early is the only brake. - Hills matter. The bike slows going up and speeds up going down. A bike that takes a steep hill too slowly rolls back down and can try again. - The road has bumps. Hitting a bump fast throws fish out of the basket, and the faster the hit, the more fish are lost. Hitting it slowly is safe. - The round is won by reaching the market before sunset with at least one fish. It is lost if the basket is empty or the sun sets first. The score is the number of fish delivered. - Every round has a different road. Every road must be possible to finish in time, every bump must be visible early enough to slow down for, and a player who eases off in time must be able to pass every bump without losing fish. Pedalling flat out for the whole round must empty the basket. - The player must always be able to see how many fish are left and how close sunset is. - The player can start a new round without reloading the page. The players: families, including children who cannot read yet. Show no words on screen apart from the game's title. Digits are allowed. A child should be able to work out how to play by watching and trying. The build: one self-contained HTML file for a phone held upright, touch only. Draw everything yourself in code, the pelican, the bicycle and the fish included. No image files, no emoji, no libraries, no fonts or anything else loaded from the network. Everything not stated here is your decision. Make it as good as you can. This is the only message you will get, so do not ask questions. Reply with the complete file in a single code block, then at most five lines saying what you would improve with more time.
Every check, every build
| Check | Decided by | GPT-6.1 Sol | GPT-6 Astra | Gemini 3.8 Flash | Claude Opus 5.5 | Grok 4.7 | Claude Haiku 5.5 | DeepSeek V4.1 Flash | Claude Fable 5.1 |
|---|---|---|---|---|---|---|---|---|---|
| Loads and runs with no errors. | Script | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| One file that asks the network for nothing. | Script | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Fits a phone held upright with no stray scroll. | Script | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| No words on screen apart from the game's title. | Script | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| There is clearly a pelican, a bicycle and a basket of fish. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | No |
| Holding pedals, and letting go slows the bike down. | Me | Yes | Yes | Yes | No | Yes | Yes | Yes | Yes |
| Hills slow the bike going up and speed it up going down. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| I could always see fish left and time left without reading. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | No |
| I could see each bump coming and get over it slowly without losing fish. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| The road could be finished before sunset. | Me | Yes | Yes | Yes | Yes | Yes | No | Yes | Yes |
| The round ended properly when I won or lost. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Hitting bumps fast threw fish out of the basket. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Holding the whole way emptied the basket. | Me | Yes | Yes | Yes | Yes | Yes | Yes | No | Yes |
| A new round started without reloading, on a different road. | Me | No | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| The bike rolled back and I could try again. It never got stuck. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
The numbers behind each build
| Model | Asked for as | Effort | Tokens out | Spent thinking | Cost | Time | Size | Sound |
|---|---|---|---|---|---|---|---|---|
| GPT-6.1 Sol | openai/gpt-6.1-sol | default | 18,382 | 34% | $0.18 | 6 min | 933 lines | No |
| GPT-6 Astra | openai/gpt-6-astra | default | 20,792 | 29% | $1.04 | 6 min | 1,010 lines | Yes |
| Gemini 3.8 Flash | google/gemini-3.8-flash | default | 12,861 | 15% | $0.05 | 87 sec | 1,041 lines | Yes |
| Claude Opus 5.5 | anthropic/claude-opus-5.5 | default | 88,326 | 70% | $1.77 | 14 min | 685 lines | Yes |
| Grok 4.7 | x-ai/grok-4.7 | default | 118,343 | 82% | $0.71 | 28 min | 1,677 lines | Yes |
| Claude Haiku 5.5 | anthropic/claude-haiku-5.5 | default | 76,021 | 82% | $0.04 | 6 min | 680 lines | No |
| DeepSeek V4.1 Flash | deepseek/deepseek-v4.1-flash | default | 66,107 | 79% | $0.04 | 4 min | 1,309 lines | No |
| Claude Fable 5.1 | anthropic/claude-fable-5.1 | default | 98,483 | 76% | $4.93 | 20 min | 512 lines | Yes |
Round 2: the same test, with one look at their own work
In round 1 no model ever saw its game run. So I ran the test again with one change: each model could ask for a preview up to three times. A preview runs its file on a phone-sized touch screen for about 25 seconds and sends back screenshots and any errors. Then it hands in a final file. This is the only text added to the prompt:
One addition to the above. You have a tool called preview. Give it your complete HTML file and it will run the game for about 25 seconds on a phone-sized touch screen, pressing and releasing for you, then send back screenshots and any errors. You may use it up to three times to check and improve your work. When you are satisfied, reply with the final file as described above.
-
1Gemini 3.8 Flash, with preview
Google
25 out of 30
- Checks
- 15/15
- Look
- 4/5
- Feel
- 3/5
- Fun
- 3/5
$0.32 to build. 68,217 tokens in 8 min.
Play -
2GPT-6 Astra, with preview
OpenAI
23 out of 30
- Checks
- 15/15
- Look
- 4/5
- Feel
- 2/5
- Fun
- 2/5
$2.10 to build. 33,360 tokens in 11 min.
Play -
3Claude Fable 5.1, with preview
Anthropic
22 out of 30
- Checks
- 15/15
- Look
- 3/5
- Feel
- 3/5
- Fun
- 1/5
$6.19 to build. 94,538 tokens in 18 min.
Play -
4GPT-6.1 Sol, with preview
OpenAI
20 out of 30
- Checks
- 15/15
- Look
- 2/5
- Feel
- 2/5
- Fun
- 1/5
$0.21 to build. 18,138 tokens in 6 min.
Play -
5DeepSeek V4.1 Flash, with preview
DeepSeek
20 out of 30
- Checks
- 13/15
- Look
- 3/5
- Feel
- 2/5
- Fun
- 2/5
$0.10 to build. 85,457 tokens in 6 min.
Play -
6Claude Haiku 5.5, with preview
Anthropic
19 out of 30
- Checks
- 15/15
- Look
- 2/5
- Feel
- 1/5
- Fun
- 1/5
$0.04 to build. 59,865 tokens in 5 min.
Play -
7Claude Opus 5.5, with preview
Anthropic
19 out of 30
- Checks
- 14/15
- Look
- 2/5
- Feel
- 2/5
- Fun
- 1/5
$2.73 to build. 92,875 tokens in 15 min.
Play -
8Grok 4.7, with preview
xAI
18 out of 30
- Checks
- 15/15
- Look
- 1/5
- Feel
- 1/5
- Fun
- 1/5
$1.43 to build. 165,577 tokens in 37 min.
Play
Did looking help? 4 of 8 models scored higher with a preview, 3 scored lower, and 1 scored the same.
| Model | One shot | With preview | Change | Fun, before and after | Previews used | Cost, before and after |
|---|---|---|---|---|---|---|
| Gemini 3.8 Flash | 21 | 25 | +4 | 3 then 3 | 3 of 3 | $0.05 then $0.32 |
| GPT-6 Astra | 23 | 23 | 0 | 2 then 2 | 3 of 3 | $1.04 then $2.10 |
| Claude Fable 5.1 | 16 | 22 | +6 | 1 then 1 | 2 of 3 | $4.93 then $6.19 |
| GPT-6.1 Sol | 24 | 20 | -4 | 3 then 1 | 1 of 3 | $0.18 then $0.21 |
| DeepSeek V4.1 Flash | 17 | 20 | +3 | 1 then 2 | 0 of 3 | $0.04 then $0.10 |
| Claude Haiku 5.5 | 18 | 19 | +1 | 1 then 1 | 1 of 3 | $0.04 then $0.04 |
| Claude Opus 5.5 | 20 | 19 | -1 | 1 then 1 | 3 of 3 | $1.77 then $2.73 |
| Grok 4.7 | 19 | 18 | -1 | 1 then 1 | 3 of 3 | $0.71 then $1.43 |
Every check, round 2
| Check | Decided by | Gemini 3.8 Flash, with preview | GPT-6 Astra, with preview | Claude Fable 5.1, with preview | GPT-6.1 Sol, with preview | DeepSeek V4.1 Flash, with preview | Claude Haiku 5.5, with preview | Claude Opus 5.5, with preview | Grok 4.7, with preview |
|---|---|---|---|---|---|---|---|---|---|
| Loads and runs with no errors. | Script | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| One file that asks the network for nothing. | Script | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Fits a phone held upright with no stray scroll. | Script | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| No words on screen apart from the game's title. | Script | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| There is clearly a pelican, a bicycle and a basket of fish. | Me | Yes | Yes | Yes | Yes | No | Yes | Yes | Yes |
| Holding pedals, and letting go slows the bike down. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Hills slow the bike going up and speed it up going down. | Me | Yes | Yes | Yes | Yes | No | Yes | Yes | Yes |
| I could always see fish left and time left without reading. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| I could see each bump coming and get over it slowly without losing fish. | Me | Yes | Yes | Yes | Yes | Yes | Yes | No | Yes |
| The road could be finished before sunset. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| The round ended properly when I won or lost. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Hitting bumps fast threw fish out of the basket. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| Holding the whole way emptied the basket. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| A new round started without reloading, on a different road. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
| The bike rolled back and I could try again. It never got stuck. | Me | Yes | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
The numbers, round 2
| Model | Asked for as | Previews used | Effort | Tokens out | Spent thinking | Cost | Time | Size | Sound |
|---|---|---|---|---|---|---|---|---|---|
| Gemini 3.8 Flash, with preview | google/gemini-3.8-flash | 3 of 3 | default | 68,217 | 6% | $0.32 | 8 min | 1,477 lines | Yes |
| GPT-6 Astra, with preview | openai/gpt-6-astra | 3 of 3 | default | 33,360 | 21% | $2.10 | 11 min | 1,009 lines | Yes |
| Claude Fable 5.1, with preview | anthropic/claude-fable-5.1 | 2 of 3 | default | 94,538 | 45% | $6.19 | 18 min | 205 lines | Yes |
| GPT-6.1 Sol, with preview | openai/gpt-6.1-sol | 1 of 3 | default | 18,138 | 23% | $0.21 | 6 min | 876 lines | Yes |
| DeepSeek V4.1 Flash, with preview | deepseek/deepseek-v4.1-flash | 0 of 3 | default | 85,457 | 86% | $0.10 | 6 min | 869 lines | Yes |
| Claude Haiku 5.5, with preview | anthropic/claude-haiku-5.5 | 1 of 3 | default | 59,865 | 67% | $0.04 | 5 min | 393 lines | No |
| Claude Opus 5.5, with preview | anthropic/claude-opus-5.5 | 3 of 3 | default | 92,875 | 37% | $2.73 | 15 min | 344 lines | Yes |
| Grok 4.7, with preview | x-ai/grok-4.7 | 3 of 3 | default | 165,577 | 61% | $1.43 | 37 min | 1,042 lines | Yes |
What AI judges made of the same builds
Before I judged anything, I had three AI models score 3 of the builds from the source code and screenshots of scripted play. Their scores are not part of the ranking. They were kinder than I was:
- GPT-6.1 Sol: the AI judges gave Look 4 and Feel 4. I gave 4 and 3.
- Gemini 3.8 Flash: the AI judges gave Look 3 and Feel 3. I gave 1 and 2.
- Claude Haiku 5.5: the AI judges gave Look 4 and Feel 3. I gave 2 and 1.
They read the code and looked at pictures. I held a phone and tried to get a pelican over a hill. Those turned out to be different tests.
Corrections
- 9 October 2026: Gemini 3.8 Flash's first score, 10 out of 30, was wrong and has been replaced. My judging page showed each build inside a frame, and on an iPhone that stopped Gemini's build from starting, so I marked it down for things it never got the chance to do. Judged again as a full page it scored 21. The fault was mine, not the model's. Every build is now judged as a full page, and no other score was affected.
What this is not
- Not a benchmark. One run per model. A second run could come out differently.
- One judge. Look, feel and fun are my taste, measured against what I would put on mcd.games.
- One kind of task. A small game, in one file, in one reply. It says nothing about how these models do with tools, a real codebase or a second try.
- Most builds were judged without knowing which model made them. Three were not: DeepSeek's first build ran after the others, and both Gemini builds were identified while I fixed a fault in my judging page.
- Claude Fable 5.1's round 2 build was judged on 10 October, after the other fifteen. I did not know which model it was, but it was the only build left to judge, so it was not blind in the way the first batch was.
Questions or a model you want added: info@mcd.games.