Live data from Hacker News

Claude Opus 5

anthropic.com

691–700 of 1001 posts

Re: Claude Opus 5

#691
post #670

Pointless anecdote: I asked it to make some slides and it decided to write its own slide rendering engine: > On the format — I dropped reveal.js and wrote a small engine inline instead. Reveal would have meant a CDN load, and a deck that half-renders because the lecture theatre wifi is flaky It one-shotted a perfect functional mini version of powerpoint (or Reveal) for a simple presentation I asked it to make.

Opus 4.8 decided to code up its own version of the SwiftUI rendering engine for iOS when I asked it to change a swipe gesture. I left the computer for several hours, came back, noticed it still wasn't done, noticed it had alarmingly burned through my weekly tokens, and had to stop it from continuing. "You're right. What I did was overkill and I should have just used iOS's built-in rendering engine. Noted for next tim…

> Noted for next time.

Is that just something it says, or will that actually affect how it will behave next time?

Re: Claude Opus 5

#692
post #104

I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code. It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is? [0] https://imgur.com/a/Nv8V7Ry

It is probably from randomness. The benchmark tasks nowadays are so long that you can't really afford to run a large number of samples of them per model & effort combination

If you're doing passes@1, especially for long-horizon agentic benchmarks, you might as well as not do the benchmark at all.

Re: Claude Opus 5

#694

Anyone else feel like Claude Code has gotten worse lately? It keeps going off on tangents I never asked about, and it won't stick to a simple rule I've given it repeatedly: stay concise, only expand when I ask. It just doesn't follow that. Worse, about two weeks ago it recommended a command and assured me it was safe. I pushed back and asked it to double-check, and it confirmed again that it was safe. I trusted that…

How do you lose weeks of work? Don’t you use git? Push it to remote? Backups?

Re: Claude Opus 5

#695
post #38

> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively. Okay so it’s worse than Opus 4.8 for my purposes I guess?

I don't think there's been any official confirmation, but even Fable safeguards seem to have gotten quite a bit of tuning, is less trigger-happy and less like regex matching.

Re: Claude Opus 5

#696

Earlier quoted context omitted.

I feel like i've seen less hype about "the next model will be agi". GPT-6 is supposed to be coming this summer, and nobody is expecting AGI now. Not sure how they're going to keep the hype cycle going

It's already happened but no one wants to admit it

We can start having this conversation when it's able to do at least 20% of the work that I have to do.

Re: Claude Opus 5

#697
post #641

My excitement about Anthropic had fabled-out dramatically when they suspended my pro account about two weeks ago within just 12 hours of fair use. I was really mind-blown when I tried Fable 5 for the first time to help me improve a game I was working on but shortly, they decided that I had a suspicious activity and suspended my account without a clear reason. I submitted a an appeal describing that I am 100% sure I h…

There’s got to be more to this story, what exactly were you up to with these models?

In the linked HN comment from a few years ago, an Anthropic employee said it was some fraud detection system.

It could be that OP is unlucky, and some of his metadata (or perhaps payment information) matches some patterns for stolen-CCs/fraud/chargebacks?

I've also found Claude.ai to be very suspicious of less mainstream browsers (e.g. Pale Moon), unfortunately.

Re: Claude Opus 5

#698
post #670

Pointless anecdote: I asked it to make some slides and it decided to write its own slide rendering engine: > On the format — I dropped reveal.js and wrote a small engine inline instead. Reveal would have meant a CDN load, and a deck that half-renders because the lecture theatre wifi is flaky It one-shotted a perfect functional mini version of powerpoint (or Reveal) for a simple presentation I asked it to make.

Opus 4.8 decided to code up its own version of the SwiftUI rendering engine for iOS when I asked it to change a swipe gesture. I left the computer for several hours, came back, noticed it still wasn't done, noticed it had alarmingly burned through my weekly tokens, and had to stop it from continuing. "You're right. What I did was overkill and I should have just used iOS's built-in rendering engine. Noted for next tim…

Lol

Re: Claude Opus 5

#699
post #446

I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from. Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move" We need an "annoying English" benchmark. - Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d... - Opus 5 Max: https…

And this is the most important observation in this thread. It’s load-bearing!

I had Fable review some legal texts yesterday.

It told me that one particular line is "the most load-bearing sentence in the document".

Fable "rated it legally load-bearing without reservation".

Re: Claude Opus 5

#700
post #635
post #382

Earlier quoted context omitted.

Here's another test of a cyberpunk ramen shop website. One thing I've found LLMs have a lot of difficulty with is angular cuts / elements that aren't easily representable with CSS. Cyberpunk aesthetics are generally a great test of that, since they have a lot of microglyphs / window decoration. Design source of truth: https://image.non.io/9d5fed20-b476-49d3-841b-37eb553fb88e.we... Opus 5 build: https://html.non.io/ne…

Just leaving this for anyone that says a design like this doesn't work: https://riceboxed.com/

[dead]
Post reply on HN