Anthropic released Claude Opus 5.5 on September 22. GPT-6 Astra has been the model I reach for day to day, so I wanted to see whether Opus could actually take that spot back, and I gave Opus 5.5, Claude Fable 5.1 and GPT-6 Astra the exact same prompts across three builds.
The builds cover illustration drawn entirely in code, a launch film with its own sound design, and a full 3D game. Every one of them is live on this site, so you can open any build below and judge it for yourself before you read what I thought.
What I tested
Opus 5.5 and Fable 5.1 ran in Claude Code and GPT-6 Astra ran in Codex, all on high effort on September 23, and each model got the prompt once and worked on its own until the build was finished and published. The launch films all pulled their images from the same image generator, so those images aren't counted in the costs below, and the costs themselves are what each model worked out from its own session logs at API list prices. Build time only counts the minutes a model was actually working.
Test 1: Illustration drawn entirely in code
Scenes from Moby Dick as a single web page printed like a risograph poster, where every sailor, wave and ink dot is drawn by the model's own code with no images at all, in the same format as Kevin Ngo's a small light.
| Opus 5.5 | Fable 5.1 | GPT-6 Astra | |
|---|---|---|---|
| Input tokens | 69.2M | 32.2M | 4.2M |
| Output tokens | 425k | 261k | 49k |
| Cost | $37.22 | $42.18 | $8.16 |
| Build time | 85 min | 63 min | 30 min |
Test 2: A launch film with its own sound design
A 20-second launch film for a product each model made up, with every sound synthesized in code, set inside a launch page that opens with a torn-paper reveal.
| Opus 5.5 | Fable 5.1 | GPT-6 Astra | |
|---|---|---|---|
| Input tokens | 26.1M | 7.5M | 4.5M |
| Output tokens | 179k | 101k | 39k |
| Cost | $11.66 | $16.55 | $7.96 |
| Build time | 40 min | 25 min | 23 min |
Test 3: 3D game development
A Tony Hawk's Pro Skater style game set in a recreation of the original Warehouse level, where every 3D model had to be built in Blender before it went into the game.
| Opus 5.5 | Fable 5.1 | GPT-6 Astra | |
|---|---|---|---|
| Input tokens | 171.3M | 23.2M | 30.7M |
| Output tokens | 520k | 161k | 150k |
| Cost | $58.05 | $32.26 | $44.08 |
| Build time | 142 min | 56 min | 49 min |
What it cost
Astra was the cheapest on two of the three builds by a wide margin, finishing Moby Dick for $8.16 while Opus 5.5 spent $37.22 and Fable 5.1 spent $42.18. The skate game went the other way, though. Fable finished it for $32.26, and Astra came in at $44.08 because it split the work across three helper agents whose tokens all count toward its total. Opus ran for more than two hours and read 171 million input tokens, most of them spent re-reading its own growing session, and that's how it got to $58.05.
| Opus 5.5 | Fable 5.1 | GPT-6 Astra | |
|---|---|---|---|
| Input tokens | 266.6M | 62.9M | 39.5M |
| Output tokens | 1.1M | 524k | 238k |
| Cost | $106.93 | $90.99 | $60.20 |
| Build time | 266 min | 143 min | 102 min |
These are API-equivalent costs: real token counts from each build, priced at each model’s published rates. The builds themselves ran on subscriptions.
Every build: cost and time
| Build | Opus 5.5 | Fable 5.1 | GPT-6 Astra | |||
|---|---|---|---|---|---|---|
| Cost | Time | Cost | Time | Cost | Time | |
| Illustration in code | $37.22 | 85 min | $42.18 | 63 min | $8.16 | 30 min |
| Launch film | $11.66 | 40 min | $16.55 | 25 min | $7.96 | 23 min |
| 3D game | $58.05 | 142 min | $32.26 | 56 min | $44.08 | 49 min |
The verdict
Opus 5.5 won all three, and it wasn't especially close. On Moby Dick it went well past what I expected, with characters and buildings like the Whaleman's Chapel and the Spouter-Inn drawn in far more detail than the other two managed. Its launch film came closest to Anthropic's own, down to small touches like the images rotating with each cut and a reverb tail under the final logo, all of it written in code. And its skate game had the most realistic character and ramps of the three, with grinds where you actually have to hold your balance the way you do in Tony Hawk's Pro Skater.
The sound design was the weak spot for all three models, which makes sense when none of them can actually listen to what they make, so that's where I'd push hardest if I ran these again.
I'll be honest, I didn't see this coming. After Astra came out I thought nothing would beat it for at least a few weeks. Astra is still what I reach for in my everyday work, the big refactors and the multi-step site architecture, because it's so methodical that I rarely have to go back and correct it. It takes longer, but it's thorough. Fable 5.1 is much faster, but it still misses things I asked it to do, and from what I've seen so far that happens a lot less with Opus 5.5.
These builds are fun, and they're not a perfect stand-in for the work you do every day, so try them yourself and decide which model you'd hand this kind of work to. If you play the skate game and have feedback, whether that's bugs or features you'd want added, I'd love to hear it, because that's the one I want to keep building out.









