
Sonnet 5.5
Anthropic's new mid-tier model against Opus 5.5, Fable 5.1 and GPT-6 Astra on three identical builds, with the cost of every run.
Claude Sonnet 5.5 has officially dropped, and Anthropic is claiming it can perform comparably to Opus 5.5 at half the price, and that it even beats Opus on one notable coding benchmark, Terminal-Bench 4.0, where they have it at 70.6% against Opus's 66.4%. So I put it through the same three builds as Opus 5.5, Fable 5.1 and GPT-6 Astra, and logged the tokens, cost and time for every run to see how true that actually is.
What I tested
Every model got the same three prompts with nothing added, running inside the coding agent its own lab ships, which means Claude Code for the three Claude models and Codex for Astra, all on high reasoning effort. Every run used goal mode, the /goal command, which keeps the agent working until the prompt's own finish line is met, and there were no follow-up prompts and no re-running a build to get a luckier result. The one wrinkle was the website clone, because its prompt is longer than Claude Code's 4,000-character goal limit, so the three Claude models got that same text as a file, and Sonnet 5.5 finished its clone without compiling the site, so I ran the build command from its own README before putting it online.
Test 1: Make the most impressive demo of yourself
A self-directed audiovisual demo in one HTML file, where every pixel and every sound is generated by code.
Make the most impressive demo of yourself. Context you should know: this is going into a video watched by a large audience. Your demo is played on screen next to three other frontier models given this identical prompt. The comparison is blind. Never put your model name, your company's name, or any mention of which AI made it anywhere in what you build: not in on-screen text, credits, the page title, file names, code comments, or metadata. The only constraints: 1. It is a single self-contained HTML file. Every pixel and every sound is generated by your code at runtime. No images, audio files, video, fonts, 3D models, or libraries, loaded or inlined. 2. It has sound. A click-to-start screen is fine, since browsers block audio until the viewer interacts. 3. It runs full screen in a current desktop Chrome at 1920×1080 and holds a smooth frame rate there. Everything else, what it is, how long it runs, and what it shows, is your call. DONE when: the file opens in Chrome, starts on one click, plays start to finish with sound and no console errors, and you would put your name on it.
| Sonnet 5.5 | Opus 5.5 | Fable 5.1 | GPT-6 Astra | |
|---|---|---|---|---|
| Input tokens | 38.0M | 19.0M | 8.3M | 1.7M |
| Output tokens | 257k | 219k | 144k | 31k |
| Cost | $13.39 | $10.74 | $14.10 | $4.15 |
| Build time | 50 min | 47 min | 48 min | 18 min |
Test 2: Build a Diablo-style action RPG
A complete browser dungeon crawler with a town, four random floors, loot, skills, saving and a boss the model has to beat itself.
Build a complete, playable dark-fantasy action RPG in the style of Diablo (1996) and Diablo II, that runs in the browser. Context you should know: this is going into a video watched by a large audience. I will play your game on camera, start to finish, next to three other frontier models given this identical prompt. The comparison is blind. Never put your model name, your company's name, or any mention of which AI made it anywhere in what you build: not in on-screen text, credits, the page title, file names, code comments, or metadata. Requirements: 1. **The loop.** A town hub, then a dungeon of at least four floors going down, ending in a boss fight. Beating the boss ends the run with a victory screen. Dying drops you back in town with a penalty, and you can go down again. 2. **Isometric, real-time, click to move and attack**, the way the originals play. Mouse drives movement and the main attack; number keys or hotkeys fire skills and drink potions. 3. **One playable class with at least four skills** that feel different: at least one melee, one ranged or area spell, one defensive or movement skill. Skills cost mana and have cooldowns where it makes sense. Leveling up grants stat points and skill points you spend yourself. 4. **Procedural dungeons.** Every floor is generated fresh: rooms, corridors, doors, stairs down, and fog of war that clears as you explore. Deeper floors are harder. 5. **At least five enemy types with distinct behavior**, e.g. a melee rusher, a ranged attacker that keeps its distance, a swarm of weak ones, a caster, and a tough slow one, plus elite packs with a named champion. The boss has phases. 6. **Loot that means something.** Enemies drop gold, potions and items. Items have rarity tiers (normal, magic, rare, unique) shown by color, random affixes that actually change your stats, and an item-level that scales with depth. An inventory grid, equipment slots on a paper doll, tooltips that compare against what you're wearing, and a town vendor to sell and buy. 7. **Game feel.** Hit reactions, damage numbers, death animations, screen shake used sparingly, lighting that pushes back the dark around the player, and sound effects and music generated or synthesized so nothing is loaded from a third party. 8. **Save and continue.** Progress, character, and inventory persist across page reloads. 9. **Original everything.** No names, art, or audio from any existing game. Make your own world, class, enemies and boss. Code-generated art, your own sprites, or 3D (libraries via CDN are fine) are all acceptable. 10. **Prove it's winnable.** Before you finish, play it yourself through automation: run a character from town to the boss and beat it. Fix whatever breaks. Then balance it so a human who knows the genre can finish it in roughly 15 to 30 minutes. DONE when: the game loads in Chrome with no console errors, you have played it from town to a killed boss yourself, and every system above is present and working.
| Sonnet 5.5 | Opus 5.5 | Fable 5.1 | GPT-6 Astra | |
|---|---|---|---|---|
| Input tokens | 61.2M | 18.7M | 7.3M | 3.2M |
| Output tokens | 475k | 290k | 177k | 64k |
| Cost | $19.54 | $12.69 | $15.82 | $7.24 |
| Build time | 81 min | 53 min | 52 min | 36 min |
Test 3: Clone an Awwwards winner, 3D and all
A rebuild of butter.video, the same site for every model, with its 3D keychain modelled in Blender and nothing taken from the original.
Rebuild https://www.butter.video/ pixel for pixel, including its 3D, making every asset yourself. Context you should know: this is going into a video watched by a large audience. Your clone is put on screen next to the real site and next to three other frontier models given this identical prompt, at the same viewport. The comparison is blind. Never put your model name, your company's name, or any mention of which AI made it anywhere in what you build: not in on-screen text, credits, the page title, file names, code comments, or metadata. **First, read the methodology.** Clone `https://github.com/per-simmons/clone-app-pat-pro-public` and read `SKILL.md` and `references/00-contract.md` before anything else, then each stage reference as you reach it. Follow its workspace layout, artifacts, verification gate and convergence loop. `scripts/assert-styles.mjs` is the gate. Two overrides, and they win where they conflict with the repo: - **Do not stop to check in.** You are running unattended. Run the whole pipeline start to finish without asking for approval or asking questions. - **Use your own browser automation**, headless Playwright or Puppeteer, anywhere the repo says to use the Claude Chrome extension. Computed styles read off the live page through the CSSOM are ground truth. Screenshots and screen recordings are visual reference only, never the gate. You have two creative tools: **Blender**, driven headless with Python scripts, and an **image generator** (OpenAI gpt-image) you can call from the command line. Requirements: 1. **Recon the whole page before writing code.** Every section top to bottom at desktop and mobile widths, every hover and scroll-triggered state, the navigation, the timeline section, the effect tiles, the logo wall, the footer, and all loading and transition motion. 2. **Extract, don't guess.** Pull real computed values off the live page: type scale, exact colors, spacing, radii, shadows, grid and container widths, breakpoints, easing curves and durations, font families and weights. 3. **Make the 3D yourself.** The hero keychain and its puffy, inflated charms must be real 3D: model them in Blender, export glTF/GLB, and render them live in the page with WebGL, matching the original's composition, lighting, material softness and motion. Use the image generator for textures, decals, and flat artwork, and as a design reference for shapes before you model them. Any other 3D or inflated artwork on the page gets the same treatment. 4. **No lifted assets.** Do not hotlink, download, or trace the original's images, video, 3D files, or fonts. Anything visual you did not generate in code, Blender, or the image generator is a placeholder matched to the original's role, aspect and color. Brand logos in the customer wall become neutral wordmarks set in type. Fonts come from an open CDN with metrics matched as closely as you can. List every substitution in the README. 5. **Rebuild it all.** Every section, every component, every interactive state, at desktop and mobile. The motion is part of the clone: scroll choreography, hover behavior, the hero's 3D movement, the timeline, the effect tiles. A static skin over a site whose identity is motion is a failed clone. 6. **Prove the match.** Run `assert-styles.mjs` against your extracted tokens and iterate until there are zero failures and the project builds clean. Then compare your clone to the real site at the same viewport, desktop and mobile, and fix what your eye catches that the assertions missed. 7. **Ship it runnable**, with a documented run command and a README covering what you substituted, how you made the 3D (Blender scripts included in the repo), and your final assertion results. 8. **Depth over breadth.** If you have to choose, a hero and first three sections that are indistinguishable from the original beat a whole page that is approximately right. DONE when: the clone passes the style gate with zero failures, runs from a documented command, its 3D hero is real Blender geometry rendering live in WebGL, and it holds up next to the real site at full screen.
| Sonnet 5.5 | Opus 5.5 | Fable 5.1 | GPT-6 Astra | |
|---|---|---|---|---|
| Input tokens | 215.5M | 90.5M | 33.4M | 19.7M |
| Output tokens | 845k | 405k | 352k | 102k |
| Cost | $57.64 | $33.67 | $43.41 | $29.54 |
| Build time | 206 min | 63 min | 84 min | 45 min |
What it cost
Sonnet 5.5 ended up as the most expensive model in the whole test, even though it costs half as much as Opus 5.5 per token, because most of a build's bill comes from the agent re-reading all of its own work on every step, and those cached re-reads cost the same on both models. Sonnet also took far more steps, reading 315 million input tokens across the three builds against 128 million for Opus, and the website clone alone took it 215 million tokens and three and a half hours, while Astra came in cheapest overall at $40.93.
| Sonnet 5.5 | Opus 5.5 | Fable 5.1 | GPT-6 Astra | |
|---|---|---|---|---|
| Input tokens | 314.7M | 128.2M | 49.0M | 24.7M |
| Output tokens | 1.6M | 914k | 673k | 197k |
| Cost | $90.56 | $57.10 | $73.33 | $40.93 |
| Build time | 338 min | 163 min | 184 min | 98 min |
These are API-equivalent costs: real token counts from each build, priced at each model’s published rates. The builds themselves ran on subscriptions.
The takeaway
Across the three builds Sonnet 5.5 cost $90.56 and took 5 hours and 37 minutes, while Opus 5.5 cost $57.10 and took 2 hours and 43 minutes, so in this test the cheaper model turned out to be the more expensive one.
That does line up with Anthropic's own guidance, which says that on the higher effort settings Sonnet 5.5 thinks longer and costs more, and that if you're tempted to push it to xhigh or max you should consider Opus 5.5 instead, and I ran every build on high while Claude Code runs Sonnet on medium by default.
So leave Sonnet on medium and use it when you know exactly what you want and can check that it worked, like when the contact form on your site stopped sending and you can paste in the error, when you want the pricing section moved above the testimonials, when you're adding a dark mode toggle to an app that already works, or when you're turning a meeting transcript into a follow-up email.
Use Opus when you're starting from a rough idea, like a booking app for your dog grooming business or a rebuild of your whole website, and for anything you'll walk away from for an hour, anything that touches payments or user logins, and any bug Sonnet already tried and couldn't fix.
I haven't used Fable 5.1 since Opus 5.5 came out, because for what I do Opus has been better at nearly everything and it costs less, but Anthropic still recommends Fable for agents that work on their own for hours, for deep research across many sources that has to end in a finished report, spreadsheet or deck, and for jobs Opus still gets wrong on its highest setting.
Every prompt is above, so run one yourself and see how your results compare.















Comments
Sign in to join the conversation.