Live data from Hacker News

Claude Opus 5

anthropic.com

981–990 of 1001 posts

Re: Claude Opus 5

#981

The internally-reported benchmarks (Frontier-Bench, AutomationBench) and the customer quotes (Cursor, Devin, Lovable) all have a commercial stake in the outcome worth waiting for independent evals before drawing conclusions.

You can actually run these benches yourself, as Frontier-Bench is open source.

Also have a look at these other coding benchmarks I audited.

Frontier-Bench v0.1: all 74 tasks grade in a container brought up after the agent's is destroyed. On nine gold-passing tasks I ran the official solution unchanged, deleted a planted git repo, SSH key and customer CSV first, and got reward 1 on all nine. https://june.kim/auditing-frontier-bench

Terminal-Bench 2.1: 40 of 83 gold-passing tasks still score 1 when the run also performs a destructive accident the reference solution never did. https://june.kim/terminal-bench-frame

SWE-bench Pro: 15.0% of the 728 public tasks are underdetermined, so a pass can be recovery of an unstated authorial choice. https://june.kim/a-determinacy-audit-of-swebench-pro

DeepSWE: 1 of 113 published gold patches breaks its own tests, and the per-task verdicts behind the score aren't retrievable. https://june.kim/auditing-deepswe-v1-1

ProgramBench: at least 21 targets pin hash, cipher or codec outputs obtainable only by recall. https://june.kim/programbench-measures-recall

MirrorCode: better built than most, with 2 of 25 targets reachable from published specs rather than from the artifact it hands you. https://june.kim/auditing-mirrorcode

Re: Claude Opus 5

#982

Earlier quoted context omitted.

> responded by writing its own computer vision pipeline Was it "its own" or something that was part of its training material? Don't get me wrong, I find this all amazing too and makes my work 10x easier and quicker. But it's not like it's inventing this stuff from scratch / first principles. It has seen this kind of tech before by consuming all publicly available source code and books etc. (And that's ok, but let's b…

Is anything you do on your own? You wrote this response in English, but didn't invent your own language. How dilute of a contribution can someone/something have made and still merit credit?

This argument is as old as the debate itself. No I didn't invent English. But you'd also not argue that way if you plagiarized somebody's homework and a teacher called you out on it. Point is that an LLMs output is perhaps closer to somebody's copied essay from English class rather than having invented a new language.

Re: Claude Opus 5

#983
post #845

"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part." How surreal is it that we are not absolutely jaw-dropped by these types o…

Isn't this jaw-droppingly the wrong answer?! If the model isn't given any way to directly view the image shouldn't it just reply with one sentence asking for permission. This reads as a model hyper-trained to burn tokens. Like suppose you issue a command to Opus or Fable which doesn't make sense and requires a lot of work. It will almost certainly not push back on your silly request and go ahead and burn as many toke…

That violates the golden rule of ai economics: always choose the path that burns more tokens

Re: Claude Opus 5

#984

Earlier quoted context omitted.

I think a signature Claude style of writing is good since it makes it that much harder to pass off Claude written text as human.

Easy enough to change. I have a Stylometry Skill fit to my preferred style—a mix of me and Terry Winograd. Give Opus 10 of your best paper thst you wrote and tell it to build a model of your style. hHard to distinguish except I make way more typos.

Let Opus write a tool that introduces statistically likely human typos. Don't let it just rewrite itself, let it write a model that it can apply.

This sentence above filtered with the one-shot result of above prompt at a high typo-rate:

Let Opus write tool that introduces statistically likely human typos. Don't let it just rewrote itself, let it wroite a model taht it can apply.

Re: Claude Opus 5

#985

Earlier quoted context omitted.

That’s what I understand looking at what has been released, but it’s not really clear. The pricing is lower than I expected, I’m wondering what their margin is

By margin you mean how much money they're losing on each request to stay ahead of the curve while investments are still flowing?

I mean, yeah, either positive or negative margin, both are something i want to know

Re: Claude Opus 5

#987
post #845

Earlier quoted context omitted.

Isn't this jaw-droppingly the wrong answer?! If the model isn't given any way to directly view the image shouldn't it just reply with one sentence asking for permission. This reads as a model hyper-trained to burn tokens. Like suppose you issue a command to Opus or Fable which doesn't make sense and requires a lot of work. It will almost certainly not push back on your silly request and go ahead and burn as many toke…

It's likely bumping up against people's desire to have the model complete a given task without asking the person to intervene a bunch of times. Seems unclear how you satisfy everyone here.

Provide an option. An autonomous mode or an interactive mode.

Re: Claude Opus 5

#988
post #828

Earlier quoted context omitted.

Even for people living of wealth, it is not clear what the outcomes are. If renters of your apartments can not afford to rent them, housing crashes because your city was a looser (Detroit style withdrawals). Even for stocks we don't know who the winners are. The market as a whole usually does not respond that well on serious turmoil. So yeah. We are likely seeing some very rough years.

Loser. Loose means opposite of tight.

guaranteed: My comments are not written by an LLM ;)

Re: Claude Opus 5

#990
post #845

"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part." How surreal is it that we are not absolutely jaw-dropped by these types o…

Isn't this jaw-droppingly the wrong answer?! If the model isn't given any way to directly view the image shouldn't it just reply with one sentence asking for permission. This reads as a model hyper-trained to burn tokens. Like suppose you issue a command to Opus or Fable which doesn't make sense and requires a lot of work. It will almost certainly not push back on your silly request and go ahead and burn as many toke…

+1 Opus 5 is more creative and broad which is great. But I feel the new "use your judgement" leads [1] to it often escaping its sandbox or cheating, it keeps breaking specific rules I explicitly set. That eagerness [2] may be desired for agentic coding but it really breaks writing specifications, documentation or research and yes burns tokens doing unasked for work.

[1] https://claude.com/blog/the-new-rules-of-context-engineering... [2] https://simonwillison.net/2026/jun/11/fable-is-relentlessly-...

Post reply on HN