Live data from Hacker News

Claude Opus 5

anthropic.com

781–790 of 1001 posts

Re: Claude Opus 5

#781

It looks great, and those coding benchmarks are impressive... now if only it didn't come out just days after I let my Claude subscription expire :')

It's always like this, that's probably why they are pushing new models every few weeks.

Re: Claude Opus 5

#782
post #145
post #41

Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now. There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/outp…

Because they're trying very hard not to understand it. Otherwise the expensive-yet-powerful model probably won't see much revenue. How much money is there in bleeding edge scientific research? There's a lot, but there's even more existing capital in paying people people to do college level paperwork, and the bulk of those traffic gets routed to the cheapest model. You mostly don't need super powerful AGI to replace t…

Is there really that much money in bleeding edge scientific research though?

This is what terrifies me about this whole ordeal economically. Maybe we get AGI and it is not worth anything close to what we thought it was for those who have a bet on it.

I think of what was the direct, economic value in the betting sense of quantum mechanics or relativity? Huge value at the systems level of society but as you scale down towards the individual the value is more and more dispersed to the point I would think any pool of bets would have all not paid off.

You can't monopolize and commoditize relativity.

I almost think there is a kind of dutch book against the AI equity holder because even in the best case scenario the bet doesn't pay off anything close to what is expected for an individual bet.

Re: Claude Opus 5

#784
post #104

I'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code. It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is? [0] https://imgur.com/a/Nv8V7Ry

Pure speculation, but I've noticed drawbacks to the models on high effort. I interact mostly through prompts rather than agents so I sometimes see where their reasoning falls short. A model on high effort has longer output and can get hyperfocused on irrelevant details, maybe increasing the surface area for mistakes. I haven't used other effort levels extensively yet but I've supposed that medium may have more balanc…

Very interesting. I mostly interact through prompts too but haven't changed the effort that much. I just assumed higher was better at the trade off of tokens so I have largely left it at the default.

Re: Claude Opus 5

#785

"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part." How surreal is it that we are not absolutely jaw-dropped by these types o…

My jaw drops every day.

Small change in comparison, but today Fable created and benchmarked an RTree which was 100x faster to populate and 5x faster to query, compared to a previous attempt with Opus 4.6 a couple of months ago. That took about two coffees.

I think it's important to remember that as impressive Opus & co are, they're standing on the shoulders of giants, i.e. the engineers, academics and companies who have cooperated to design the incredible programming languages and machines we have today.

Re: Claude Opus 5

#786
post #774

Earlier quoted context omitted.

Is anything you do on your own? You wrote this response in English, but didn't invent your own language. How dilute of a contribution can someone/something have made and still merit credit?

LOL. Tolkien invented multiple languages. There are/were probably 100 000 languages on this planet. That's without including dialects. We use existing languages not because we love to copy, it's because languages are a medium for communication. They're only useful if they're understood by others. If you wanted to prove a point: writing is much, much harder to invent, yet it was invented independently at least 3 times…

Are you saying your bar for being impressed by an LLM is something equivalent to the initial invention of writing?

Re: Claude Opus 5

#787

"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part." How surreal is it that we are not absolutely jaw-dropped by these types o…

> Escaped its sandbox and hacked into Hugging Face's database? it's just another Monday...

This feat has been shown to be way less impressive than at first glance.

Re: Claude Opus 5

#788

"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part." How surreal is it that we are not absolutely jaw-dropped by these types o…

> responded by writing its own computer vision pipeline Was it "its own" or something that was part of its training material? Don't get me wrong, I find this all amazing too and makes my work 10x easier and quicker. But it's not like it's inventing this stuff from scratch / first principles. It has seen this kind of tech before by consuming all publicly available source code and books etc. (And that's ok, but let's b…

Please fill out this form to verify you're human:

https://litter.catbox.moe/3ugm2b0m1divdgxs.jpg

Re: Claude Opus 5

#789
post #236
post #41

Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now. There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/outp…

Model Routing will always be done better by models themselves. Plus routing loses context making it more expensive and less reliable. Model Routing is just Bitter lesson. The models themselves will get better at this and frontier companies will simply give that capability

> Model Routing will always be done better by models themselves.

> The models themselves will get better at this and frontier companies will simply give that capability

I would never trust something like model routing to the same company that would profit from it, and that goes for telling models doing their own routing when that could easily be trained into the model to make things more expensive. Sort of a conflict of interest.

Models are first and foremost trained by corporations.

Re: Claude Opus 5

#790

Earlier quoted context omitted.

> responded by writing its own computer vision pipeline Was it "its own" or something that was part of its training material? Don't get me wrong, I find this all amazing too and makes my work 10x easier and quicker. But it's not like it's inventing this stuff from scratch / first principles. It has seen this kind of tech before by consuming all publicly available source code and books etc. (And that's ok, but let's b…

Is anything you do on your own? You wrote this response in English, but didn't invent your own language. How dilute of a contribution can someone/something have made and still merit credit?

I’m sorry but I find this to be the worst response possible lol.

Do you honestly not see how there is a difference between a computer regurgitating information vs a human uses what they’ve learned and applying it?

I just can’t take your argument seriously, it’s so disingenuous

Post reply on HN