simonw
6 hours ago
Pelicans riding bicycles for Haiku at the different thinking levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right.
The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds.
The most recent release of my llm-anthropic plugin queries the Anthropic model listing API directly, so I didn't have to upgrade the plugin to add support for this model:
llm install llm-anthropic -U
llm anthropic refresh
llm -m claude-haiku-5.5 'prompt goes here'
EDIT: Here's the Haiku 4.5 pelican from a year ago for comparison, it was terrible: https://simonwillison.net/2025/Oct/15/claude-haiku-45/rotis
2 hours ago
thinking_effort: max Reasoning trace: This is the classic pelican-on-bicycle SVG test.
Opus 5.5 had similar response on max: This is a classic test request
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
I think your test is already embedded into the models. You should search for new frontier tests to subject the models to. Maybe they should now try to unify the standard model and general relativity in physics. I'm pretty sure this is nowhere to be found in any training data nor shared in any chat between a scientist and a LLM ;)
simonw
2 hours ago
The fact that they've heard of the test doesn't seem to help them draw a good picture of a pelican riding a bicycle.
That said... here's "Generate an SVG of an armadillo in fishnet tights jaywalking on Mars" on xhigh for comparison: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... (and here's the same thing from other models: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...)
monster_truck
an hour ago
Maybe they're as sick of it as some of us are?
breezybottom
an hour ago
LLMs hallucinate and pander. Calling something "classic" is a common LLMism
ozgung
5 hours ago
Of course, the sun again. Everyone knows that a pelican can't ride a bicycle without a sun in the frame and can only go right.
accrual
4 hours ago
I wonder if we'll start to see pelicans like a mascot of sorts. You could have a pelican pin on your backpack.
> "What's up with the pelican?"
Well you see in the early days of LLMs we wanted a fun way to test new models, and there was this blog, ...
snowram
4 hours ago
Will Smith eating spaghetti is the OG benchmark
SwellJoe
an hour ago
Do people really want to be seen making using AI part of their personality?
ultrarunner
41 minutes ago
Several people I work with have adopted its mannerisms unironically. Then again we don't exactly hang out outside of work.
kenhwang
2 hours ago
The grass too, everyone knows bikes ride on grass.
cortesoft
3 hours ago
The medium thinking effort one doesn't have a sun at all?
Jcampuzano2
4 hours ago
I always find the time/token differences between the xhigh and the max effort levels for Claude models absolutely insane.
Even more so, because in a lot of their benchmarks they use the max models. I honestly think I'd rather these labs use their xhigh models as the default for benchmarking instead since I don't think the average person is even using max.
mudkipdev
3 hours ago
Benchmarks are the entire reason why max exists
LoganDark
3 hours ago
I use max all the time, a bit annoyed that they keep trying to silently switch me off it. (Claude Code will refuse to remember a setting of max and will continually reset it to xhigh - I have an objection to these patterns in general)
I'm definitely not the average person though.
abustamam
3 hours ago
I say half facetiously - have you tried writing a skill or rule to remember your setting as a workaround?
I actually don't like that it sometimes remembers the last model/effort i used. I should be able to set a default model/effort that is separate from the one off fable runs I use.
LoganDark
3 hours ago
I thought the thinking effort was specified out of band from that, though maybe it's not. Not sure if the model was trained to listen in other areas. The biggest issue is, it's difficult to tell if it works because you can no longer see the thinking! Though I guess if you can't tell a difference in the output, was there any point to max in the first place?
internet101010
an hour ago
Long legs to illustrate me attempting to stretch my budget by using Haiku the last week of the month.
zahlman
3 hours ago
I'm still getting network errors. Seems to be CORS-related.
ijidak
5 hours ago
I find it helpful when you post your link that compares the model to other models in the same class or family, or shows progression over time.
The pelicans all start to look the same after a while.
But seeing the comparison to other models by class, family, or historical progression gives an excellent frame of reference.
simonw
4 hours ago
Good call, I've edited my comment.
Here's the Haiku 4.5 pelican from a year ago - it sucked in comparison to Haiku 5.5: https://simonwillison.net/2025/Oct/15/claude-haiku-45/
rattray
3 hours ago
How does Haiku 5.5 compare with modern alts in its class, like Luna-6 or OSS models of similar speed/cost?
simonw
2 hours ago
You can browse through other models here - Haiku is doing pretty well now compared to Luna et al: https://simonwillison.net/tags/pelican-riding-a-bicycle/
rattray
an hour ago
Hmm… is there like a chart that shows rough quality vs cost?
vinni2
5 hours ago
I thought Anthropic models didn’t generate images.
simonw
5 hours ago
This is SVG, but recent Anthropic models have got extremely good at other forms of visual data.
Here's a Blender model I had Claude Opus 5.5 create: https://tools.simonwillison.net/blender-viewer?url=https%3A%...
And here's some animated pixel art by Opus 5.5: https://tools.simonwillison.net/kakapo-party
And some Monkey Island style music (Opus can compose music too): https://tools.simonwillison.net/scrimshaw-jukebox
Anthropic's models do all of this by outputting code. GPT-6 Astra has similar capabilities - I got this Blender model using that: https://tools.simonwillison.net/blender-viewer?url=https%3A%...
noduerme
3 hours ago
Pardon, I have a lot of questions about that Scrimshaw music text format. It's clever. Did you invent it, and is it specifically intended to be written to by LLMs? Is the editor/player LLM-coded as well, and was this its recommendation for a format that would be easy for LLMs to write? I'm wondering why this instead of say, asking it to write a .MOD file.
simonw
2 hours ago
Opus invented it, and wrote the player, and the songs.
My prompts were:
> I want you to write some computer game music for me. First design simple text based format for the music and build an artifact that can play it out loud - include some example tracks in that artifact
> I am looking for music of the quality of the original secret of Monkey Island
And then later:
> Modify scrimshaw jukebox to add a copy-paste prompt that explains the music format, it should be shown at the bottom of the page below the readable instructions, the prompt should be designed to help any LLM tool compose music in the correct format. It should have a copy to clipboard button.
https://claude.ai/share/1f721c20-2499-4d23-b368-3ab57146d956 and then https://claude.ai/code/session_01R3xuRtjVHqHo1GNbget6Tu
VectorLock
an hour ago
What was the prompt for the widow's bay egg?
simonw
an hour ago
I pasted in a concept image (created by ChatGPT images here https://chatgpt.com/share/6ac16336-c288-83e8-9f28-084d184363...) and prompted:
> Use your blender local skill to create a blender model of this faverge egg
The blender local skill is this one: https://github.com/simonw/gpt-6-astra-blender-pelican-bicycl... - which I described here: https://til.simonwillison.net/llms/blender-coding-agents-mac...
lastdong
4 hours ago
This is great! Love the pixel art and tunes.
jansan
5 hours ago
They are really good at generating artifacts, which are windows within the replies containing all kind of visualization, often interactive.
They are still not great at SVG. I just asked Opus and Fable to add a background to an SVG and the results were, well, not great.
mayli
5 hours ago
SVG is hard.
hazelnut
5 hours ago
Tried it with GPT-6 Astra with Ultra but the outcome was underwhelming with Blender. Maybe it was my prompting ¯\_(ツ)_/¯
vunderba
5 hours ago
I’ve been playing around with Opus 5.5 which has made a big leap over previous generations in its ability to use a simple drawing-instruction prompt to generate images.
This creates Sierra AGI-style adventure game scenes painted live from simple Turtle-esque drawing instructions so you can basically provide it an empty canvas and then position text labels on the canvas where you want certain things (tavern, oak tree, etc) and it will generate a custom script for rendering them in a EGA graphics style.
abirch
5 hours ago
They generate svg. You can paste in pngs and they'll convert them to svg with varying degrees of success.
sixtyj
5 hours ago
They don’t do raster images.
1ucky
5 hours ago
Those are SVGs not images.
jonshariat
5 hours ago
SVG is code
fakedang
5 hours ago
They're SVGs
sixothree
5 hours ago
I've created multiple videos using Claude Code, including music and speech. It generates python which in turn generates frame PNGs that it runs through ffmpeg.
Please don't judge me too harshly for this particular poop video. But here is an example of something 100% generated with claude prompts only.
zyberzero
4 hours ago
To clarify the ”100%” part - the Python script generated the video output, and you did nothing? No video edit at all? Then I think it is impressive! Are you able to share the prompts you used?
sixothree
4 hours ago
Source code is linked from the video! Scan the QR code. I will try to /resume tonight and give you some prompts.
cruffle_duffle
3 hours ago
I mean Claude code sessions are all jsonl files that it can interrogate on its own. Get a new agent to capture how it was made and what the prompts were. No need for tedious /resume’ing and prompting.
sixothree
an hour ago
It's the easiest way to convey what I was doing. Anyway, some prompts added. I have other videos too if intersted.
sixothree
an hour ago
I'm going to truncate a lot of this. But here is the root prompt that will get you a video.
Make a youtube poop video about having too many tabs that keep appearing faster than you can close them. Style it after the "too many cooks" viral youtube video that seemed to repeat over and over, getting worse and worse every time. You have ffmpeg and a plethora of programming languages at your disposal. Go completely nuts and make it as extensive and creative as you want. Render should be 1920x1080, and later 4k if you did a good job.
[discussions about acts, characters, darkening theme, etc.]Workshopping the acts:
The play has a font issue. See screenshot.
[Image #1]
Also, can you change the thing that happens at the start of episode one the cursor clicking the one tab and it becomes two. Then it clicks another tab and it becomes four. Is that a breaking change? You can change the speed of the actions to fit it in if needed.
Here is a refinement prompt: The last screen just before "The End", the cursor clicks on the browser window instead of the tab close button. Can you make it click on the tab itself? Here is a screenshot of where it clicks [Image #3]
Also, on the intro screen, the cursor clicks the tab and new tabs open, can you have it click the links in the page instead. See screenshot[Image #4].
You can keep the exact same timing for both of these.Bluestein
4 hours ago
The Purple Screen of Death at the end :)
ghoshbishakh
5 hours ago
Bruh. Svg. It is like drawing something with geometric shapes which are represented using equations.