$ tail -f stew
OpenAI's latest models are out, and the pitch is bold: frontier-level performance that beats Anthropic at a much friendlier price. So I put that claim to the test with my favourite benchmark: getting each model to draw an SVG of a hedgehog playing a violin. Drawing an SVG is much harder for the models than using normal generative AI.
Some were adorable. Some were... roadkill.
By the end, one lab clearly delivers more quality per prompt and suggest where you should spend your money when only the best result will do.
See every hedgehog in full glory (and horror) the side-by-side →
The fight has moved to price and tokens-per-answer instead of headline specs. Sonnet 5 and Terra land almost on top of each other on output price. Fable 5 is the most expensive token on the board by a wide margin.
Full price and performance breakdown, plus a cost calculator that re-sorts all six models by your own token budget: price & performance →