Live data from Hacker News

Claude Opus 5

anthropic.com

651–660 of 1001 posts

Re: Claude Opus 5

#651

I wonder if this is one of the few times simonw's pelican was broken on the first try [1]: https://tools.simonwillison.net/markdown-svg-renderer#url=ht... My experience with Opus 5 thus far haven't been that great either. It's been making mistake after mistake editing my coding plans that were being reviewed by GPT-6 Sol. [1] https://simonwillison.net/2026/Jul/24/introducing-claude-opu...

Yeah this was pretty surprising. For almost every model I've run the pelican against the first attempt was at least recognizable enough that I didn't feel like the model needed a second shot.

It's always a roll of a dice, but it's surprising that the dice rolls so infrequently come up bad, yet Opus 5 rolled a bad pelican this one time.

I suspect it's just a freak occurrence. I rolled a few more and they were all fine, I think Opus 5 just got unlucky.

Re: Claude Opus 5

#652
post #382

Earlier quoted context omitted.

Here's another test of a cyberpunk ramen shop website. One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration. Design source of truth: https://image.non.io/9d5fed20-b476-49d3-841b-37eb553fb88e.we... Opus 5 build: https://html.non.io/ne…

God damn, we are living in the future. I love this so much. Designs like this would never have seen the light of day in the cellphone incrementalism / corporate memphis era of tech. Now people can be weird and awesome again. This is 1980's cyberpunk / late-90's Matrix / early-00's sci-fi UI. Great ideas that died to frutiger aero (which isn't a bad design aesthetic) and flat design (which is). This is fun and it's go…

Youre delusional fella.

The reality is the web is going to turn into a walled garden.

Re: Claude Opus 5

#653
post #382
post #332

Doing testing with it now, specifically for image->html conversion. Previously Fable was the best at this, followed by Gemini 3.1 pro (a surprising #2, but Google has great vision models). Opus' results seem to be more accurate than Fable, following the design source of truth better. Example results: Design source of truth: https://image.non.io/73e239a3-880f-4793-b65f-4810be2d9378.we... Opus 5 build: https://html.non…

Here's another test of a cyberpunk ramen shop website. One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration. Design source of truth: https://image.non.io/9d5fed20-b476-49d3-841b-37eb553fb88e.we... Opus 5 build: https://html.non.io/ne…

The last 10% is gonna take 90% of the time though.

Re: Claude Opus 5

#654
post #446

I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from. Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move" We need an "annoying English" benchmark. - Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d... - Opus 5 Max: https…

The user is making an extremely sharp point – full stop.

Re: Claude Opus 5

#655

What's the point of 150 pages description of a model that's going to be replaced in a couple months? Who even reads this? I know it's cheap to generate text with LLMs, but this is just noise at this point.

I actually do read them. Not in severe detail, but not casually either. 150 pages is really not very long and there doesn't seem to be too much bloat. (I would cut out the moral personhood stuff but that's a political/ideological thing). This is snarky but I am grumpy: I wonder if there's a correlation between me refusing to use LLMs and me being happy to read a novella-sized PDF about them.

"I wonder if there's a correlation between me refusing to use LLMs and me being happy to read a novella-sized PDF about them."

Well, one possible explanation is that you have time on your hands.

Lots of people using LLMs do so because they are in a hurry to do or ship something. In your case it would appear you have time budget for reading.

Re: Claude Opus 5

#656
post #367

Earlier quoted context omitted.

I don't understand how the data retention works. My company has an enterprise license with no data retention but if I ask Claude about past conversations, it remembers. So surely the information is being stored somewhere

Opus 4.7+ and Fable are both much more aggressive than prior models with respect to writing memories to a location that's effectively quasi-private for them. It's device-local (so passes retention constraint), and you can see it, but only if you go looking for it. It's a funny design/affordance. I do see them often writing memories of things that that feel unlikely to be important going foward / with other tasks, but…

Another silent inflation of token count. There's no force on earth that can overcome that incentive for the labs.

Re: Claude Opus 5

#658
One of the best hamsters [0].

Again, their "none" version costs more than "low", and says zero reasoning tokens, makes no sense[1].

As always, the "low" version seems to be the best price/perf ratio for factual answers and tool usage, and high one for creative tasks (coding, generating UIs, etc.)

[0]: https://aibenchy.com/compare/anthropic-claude-opus-5-high/an...

[1]: https://aibenchy.com/compare/anthropic-claude-opus-5-high/an...

Comparison with other top models (5.6 Sol, 3.6 Flash, Kimi K3): https://aibenchy.com/compare/anthropic-claude-opus-5-high/op...

Re: Claude Opus 5

#659
post #483

Earlier quoted context omitted.

Fable might be using those phrases less, but its writing is still terrible and exhausting to read.

Could these complex/hard to read Fable outputs be sign of some kind of industrial level of intelligence, which us humans may have a hard to comprehend, while it may be also hard for machine to use simpler texts to properly outline all nuances and complexities of concepts it output?

It's more like Claude models entirely suck at extracting key points. No matter how hard I emphasize that it needs to pick the "load-bearing" facts and claims, it cannot stop itself muttering around. It never nails the core logical structure. GPT is better at that.

Re: Claude Opus 5

#660
post #467
post #457

It really feels as though my 20 year career as a front end developer is coming to a very abrupt end; at least as I have know it these past two decades.

really? I have yet to see fable or 5.6 reliably generate front end code with correct a11y, for one thing -- does that not matter to the work you do?

It does, when you explicitly ask you to do it. Hence the need for experienced devs running those models.
Post reply on HN