MCDGames

Version 1: one try per picture, least thinking asked for, and 12 models thought anyway. Version 2 adds repeat tries, a thinking-hard set, a judge's score and a 100 shape check.

CatTreeHouseSnug blockAxolotl

Tap any picture to see it bigger. Each one is the model's first and only try.

Some models think before they answer and some do not. I asked every model for as little thinking as it allows, so this is close to a quick sketch from each. The small note beside each name says how much it thought anyway.

Four fuzzy pastel blocks with small smiling faces, stacked: pink, blue, green and purple

These are real snug blocks, from my game SnugBlox. The models never saw them. They were only told to draw: "a soft pink square with rounded corners, two black dot eyes and a small smile".

How it was done
  • Each picture is one message to the model and one reply. No other instructions, no tools, and no second try.
  • The model does not make an image. It writes a list of up to 40 shapes with positions and colours, and this page draws them. Nothing the model wrote is run as code.
  • A picture that could not be read at all shows as "no picture".
  • Thinking was turned as low as each model allows. That is not the same for all of them: the oldest models cannot think at all, some can switch it off, and some think a lot whatever they are asked. Average thinking per picture, in tokens (a token is about three quarters of a word): GPT-3.5 Turbo 0, GPT-4 0, GPT-4o 0, GPT-5 0, GPT-5.5 47, GPT-6.1 Sol 0, Claude Sonnet 4 0, Claude Sonnet 4.5 0, Claude Sonnet 4.6 0, Claude Sonnet 5 0, Claude Sonnet 5.5 0, Gemini 2.5 Flash 715, Gemini 3 Flash 0, Gemini 3.5 Flash 0, Gemini 3.8 Flash 0, Claude Opus 4.1 0, Claude Haiku 4.5 0, Claude Opus 4.5 0, Claude Opus 5 0, Claude Fable 5.1 0, Claude Opus 5.5 0, Claude Haiku 5.5 0, Kimi K2 0, Kimi K2.5 0, Kimi K2.6 5624, Kimi K3 169, Qwen 2.5 72B 0, Qwen3 235B 0, Qwen3.5 Plus 5938, Qwen3.7 Max 1845, Qwen3.8 Max Prime 1193, DeepSeek V3 0, DeepSeek V3.1 0, DeepSeek V3.2 0, DeepSeek V4 Flash 1200, DeepSeek V4.1 Flash 0, Grok 4.20 6405, Grok 4.5 321, Grok 4.7 612, Mistral Large 0, Mistral Large 2 0, Mistral Large 3 0, Mistral Large 4 3624, Llama 3.1 70B 0, Llama 3.3 70B 0, Llama 4 Maverick 0.
  • More thinking would probably give better pictures from the models that have it. That has not been tested here.
  • Models are reached through OpenRouter, on their maker's own servers where the maker still runs them there. These ran somewhere else: Claude Sonnet 4 (Amazon Bedrock), Claude Opus 4.1 (Amazon Bedrock), Kimi K2 (Novita), Kimi K2.5 (Amazon Bedrock), Qwen 2.5 72B (DeepInfra), Qwen3 235B (DeepInfra), DeepSeek V3 (DeepInfra), DeepSeek V3.1 (CoreWeave), DeepSeek V3.2 (AtlasCloud), DeepSeek V4 Flash (Alibaba), Llama 3.1 70B (Amazon Bedrock), Llama 3.3 70B (AkashML), Llama 4 Maverick (DigitalOcean).
  • The models: GPT-3.5 Turbo (openai/gpt-3.5-turbo, released 2023-05-27); GPT-4 (openai/gpt-4, released 2023-05-27); GPT-4o (openai/gpt-4o, released 2024-05-12); GPT-5 (openai/gpt-5, released 2025-08-07); GPT-5.5 (openai/gpt-5.5, released 2026-04-24); GPT-6.1 Sol (openai/gpt-6.1-sol, released 2026-09-29); Claude Sonnet 4 (anthropic/claude-sonnet-4, released 2025-05-22); Claude Sonnet 4.5 (anthropic/claude-sonnet-4.5, released 2025-09-29); Claude Sonnet 4.6 (anthropic/claude-sonnet-4.6, released 2026-02-17); Claude Sonnet 5 (anthropic/claude-sonnet-5, released 2026-06-30); Claude Sonnet 5.5 (anthropic/claude-sonnet-5.5, released 2026-09-28); Gemini 2.5 Flash (google/gemini-2.5-flash, released 2025-06-17); Gemini 3 Flash (google/gemini-3-flash-preview, released 2025-12-17); Gemini 3.5 Flash (google/gemini-3.5-flash, released 2026-05-19); Gemini 3.8 Flash (google/gemini-3.8-flash, released 2026-09-02); Claude Opus 4.1 (anthropic/claude-opus-4.1, released 2025-08-05); Claude Haiku 4.5 (anthropic/claude-haiku-4.5, released 2025-10-15); Claude Opus 4.5 (anthropic/claude-opus-4.5, released 2025-11-24); Claude Opus 5 (anthropic/claude-opus-5, released 2026-07-24); Claude Fable 5.1 (anthropic/claude-fable-5.1, released 2026-09-01); Claude Opus 5.5 (anthropic/claude-opus-5.5, released 2026-09-22); Claude Haiku 5.5 (anthropic/claude-haiku-5.5, released 2026-10-07); Kimi K2 (moonshotai/kimi-k2, released 2025-07-11); Kimi K2.5 (moonshotai/kimi-k2.5, released 2026-01-26); Kimi K2.6 (moonshotai/kimi-k2.6, released 2026-04-20); Kimi K3 (moonshotai/kimi-k3, released 2026-07-16); Qwen 2.5 72B (qwen/qwen-2.5-72b-instruct, released 2024-09-18); Qwen3 235B (qwen/qwen3-235b-a22b-2507, released 2025-07-21); Qwen3.5 Plus (qwen/qwen3.5-plus-02-15, released 2026-02-16); Qwen3.7 Max (qwen/qwen3.7-max, released 2026-05-21); Qwen3.8 Max Prime (qwen/qwen3.8-max-prime, released 2026-09-23); DeepSeek V3 (deepseek/deepseek-chat, released 2024-12-26); DeepSeek V3.1 (deepseek/deepseek-chat-v3.1, released 2025-08-21); DeepSeek V3.2 (deepseek/deepseek-v3.2, released 2025-12-01); DeepSeek V4 Flash (deepseek/deepseek-v4-flash, released 2026-04-23); DeepSeek V4.1 Flash (deepseek/deepseek-v4.1-flash, released 2026-09-10); Grok 4.20 (x-ai/grok-4.20, released 2026-03-31); Grok 4.5 (x-ai/grok-4.5, released 2026-07-08); Grok 4.7 (x-ai/grok-4.7, released 2026-09-21); Mistral Large (mistralai/mistral-large, released 2024-02-25); Mistral Large 2 (mistralai/mistral-large-2407, released 2024-11-18); Mistral Large 3 (mistralai/mistral-large-2512, released 2025-12-01); Mistral Large 4 (mistralai/mistral-large-4-0, released 2026-10-06); Llama 3.1 70B (meta-llama/llama-3.1-70b-instruct, released 2024-07-22); Llama 3.3 70B (meta-llama/llama-3.3-70b-instruct, released 2024-12-06); Llama 4 Maverick (meta-llama/llama-4-maverick, released 2025-04-05).
  • Everything, including an earlier test of two drawing styles, cost $4.34.

The message a model gets

Draw {thing} using simple shapes.

The canvas is 100 wide and 100 tall. The point 0,0 is the top left corner. Reply with a JSON list of shapes. Shapes are drawn in order, so later shapes sit on top of earlier ones. Each shape is one of these five kinds:
{"shape":"circle","x":50,"y":50,"r":10,"color":"#3366cc"}
{"shape":"ellipse","x":50,"y":50,"rx":20,"ry":10,"color":"#cc9933"}
{"shape":"rect","x":10,"y":10,"w":30,"h":20,"color":"#339966"}
{"shape":"polygon","points":[[10,10],[50,90],[90,10]],"color":"#993399"}
{"shape":"line","x1":10,"y1":10,"x2":90,"y2":90,"width":2,"color":"#333333"}

Use at most 40 shapes. Reply with the JSON list and nothing else.

Where it says {thing}: "a cat"; "a tree"; "a house"; "a snug block: a soft pink square with rounded corners, two black dot eyes and a small smile"; "a cute pink axolotl plush toy with big eyes and lighter pink gills".

What this does not show
  • One try per picture. A model can do better or worse on another day.
  • This is drawing by writing down shapes, which is a strange way to draw. It is not the same as an image generator.
  • The year is when the model was first listed on OpenRouter, which is close to when it came out.
  • Newer does not mean better at everything. It is 5 pictures.

All the tests