Live data from Hacker News

Claude Opus 5

anthropic.com

641–650 of 1001 posts

Re: Claude Opus 5

#641

My excitement about Anthropic had fabled-out dramatically when they suspended my pro account about two weeks ago within just 12 hours of fair use. I was really mind-blown when I tried Fable 5 for the first time to help me improve a game I was working on but shortly, they decided that I had a suspicious activity and suspended my account without a clear reason. I submitted a an appeal describing that I am 100% sure I h…

There’s got to be more to this story, what exactly were you up to with these models?

Re: Claude Opus 5

#642
post #177

Anyone has an insight into how much money labs are putting into benchmarks? Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!). Likely just a drop in the bucket to the training budget but still..

It's $0/close to 0, they aren't at 100% demand so any leftover compute is "not spent".

The cost they are quoting is API cost, so it's already inflated on that.

Re: Claude Opus 5

#643

My excitement about Anthropic had fabled-out dramatically when they suspended my pro account about two weeks ago within just 12 hours of fair use. I was really mind-blown when I tried Fable 5 for the first time to help me improve a game I was working on but shortly, they decided that I had a suspicious activity and suspended my account without a clear reason. I submitted a an appeal describing that I am 100% sure I h…

[flagged]

It doesn’t read as AI to me. The grammar is human-level quality.

Re: Claude Opus 5

#644

I've yet to understand why they call a 190 page PDF a "card". Calling something a card invokes a small, quick rundown of pertinent details, not every single possible detail.

Because "model card" is a set phrase, it's a concept. It originates from a time when they were shorter. Like datasheets, even if it's not literally a sheet. They could say "tech report" but model card makes it clear that it's a specific kind of tech report.

I think 'model card' should be a 1 page summary of key info. The report format should be something like a 'model data sheet' (like safety data sheets that you get with chemicals). 98% of people would only want to know the key info, not read a whole report.

Re: Claude Opus 5

#645
post #382

Earlier quoted context omitted.

Here's another test of a cyberpunk ramen shop website. One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration. Design source of truth: https://image.non.io/9d5fed20-b476-49d3-841b-37eb553fb88e.we... Opus 5 build: https://html.non.io/ne…

I really like that design. May I ask the name of the website builder/diagram? Is it Relume?

This is my own tool, diffui. Thanks, though I will say I spent like... 7 minutes on this. Feel free to completely lift the design or the implementation.

Re: Claude Opus 5

#646
post #483

Earlier quoted context omitted.

Fable might be using those phrases less, but its writing is still terrible and exhausting to read.

Could these complex/hard to read Fable outputs be sign of some kind of industrial level of intelligence, which us humans may have a hard to comprehend, while it may be also hard for machine to use simpler texts to properly outline all nuances and complexities of concepts it output?

I’m not convinced. In humans intelligence often means someone is better at explaining and needs fewer words to do so.

Re: Claude Opus 5

#647

I had a moderately complex review in a large C/C++ codebase that Codex/GPT-5.6-sol already cleaned up so I threw it at Opus 5. 4 errors found. That seemed odd, so I handed it back to GPT. All were false. Opus doesn't seem to look at the wider context and understand which functions were called in certain contexts. I gave GPT's analysis back to Opus and it admitted its mistake. Maybe it's good for writing code, but as…

A bit worrying that at no point in here did you say you investigated the errors. LLM 1 is disagreeing with LLM 2. Shouldn’t you be the tie breaker?

Re: Claude Opus 5

#648
post #97
post #17

https://www.anthropic.com/news/claude-opus-5 - A blog post for those not wanting to go through a 190ish page pdf

I like how they highlighted Opus 5 as the best for “Agentic Coding” even though the number is slightly lower than Fable. Close enough for marketing, I guess!

And no data retention for 30 days.

Re: Claude Opus 5

#649
post #635
post #382

Earlier quoted context omitted.

Here's another test of a cyberpunk ramen shop website. One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration. Design source of truth: https://image.non.io/9d5fed20-b476-49d3-841b-37eb553fb88e.we... Opus 5 build: https://html.non.io/ne…

Just leaving this for anyone that says a design like this doesn't work: https://riceboxed.com/

This is a good version of that design style though

Re: Claude Opus 5

#650
I wonder if this is one of the few times simonw's pelican was broken on the first try [1]:

https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

My experience with Opus 5 thus far haven't been that great either. It's been making mistake after mistake editing my coding plans that were being reviewed by GPT-6 Sol.

[1] https://simonwillison.net/2026/Jul/24/introducing-claude-opu...

Post reply on HN