Live data from Hacker News

Claude Opus 5

anthropic.com

961–970 of 1001 posts

Re: Claude Opus 5

#961

Earlier quoted context omitted.

There's been a dramatic increase in sloppy kinda-working something.vercel.com output in niche circles. I think that shows at least a weak demand for certain types of software, but it's not clear that the vast piles of cash that have been burned producing these "apps" are actually resulting in quality of life improvements.

I’m waiting for like GEICO’s app to significantly improve. I’m not expecting Chrome to because it already had virtually unlimited resources devoted to it. On the other side I’m not expecting my local water district’s to any time soon because the people in charge don’t care even if it was significantly cheaper to improve. But in the middle is something like GEICO and that’s where I expect improvements if this thing is…

Except GEICO probably believes it gets no more money from improving the app. Instead they will attempt to improve their general risk modeling as that's where they think the money will come from.

Re: Claude Opus 5

#962
post #691

Earlier quoted context omitted.

Opus 4.8 decided to code up its own version of the SwiftUI rendering engine for iOS when I asked it to change a swipe gesture. I left the computer for several hours, came back, noticed it still wasn't done, noticed it had alarmingly burned through my weekly tokens, and had to stop it from continuing. "You're right. What I did was overkill and I should have just used iOS's built-in rendering engine. Noted for next tim…

> Noted for next time. Is that just something it says, or will that actually affect how it will behave next time?

My experience is that unless it actually writes it down somewhere, it's almost definitely going to forget within the next time or two that the context gets compacted. Even if it does write a rule down, it still may or may not actually follow the rule because it just becomes one additional piece of context to weigh alongside everything else.

To be clear, none of this is specific to Claude as much as just a property of how these tools work. Certain models might be better or worse at making the right calls here, but the fundamental constraint of context limiting how much knowledge can be retained and a "rule" in context, whether written down in advance or manually remembered for the time being, is not a guarantee it will be followed.

Re: Claude Opus 5

#963

"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part." How surreal is it that we are not absolutely jaw-dropped by these types o…

ChatGPT released in 2022. We been promised to reach AGI in 2024, then 2025, then 2026, then 2028. It is 2026 and we still do not have AGI. In 2028 we will have jump to 40% in ARC-AGI 4, and you will say "Oh it's been less than 6 years since ChatGpt thing".

Re: Claude Opus 5

#964

Why should I be amazed at something that promises to destroy my life? I genuinely don't understand why people who have to work for their living are amazed at this. It will have a vast negative impact on your life unless you already live off of your wealth.

Could you please stop posting in the flamewar style and/or taking HN threads on generic-indignant tangents? Your account has been doing it repeatedly, and we're trying for something different here.

If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.

We detached this subthread from https://news.ycombinator.com/item?id=49045107.

Re: Claude Opus 5

#966
post #236

Earlier quoted context omitted.

Model Routing will always be done better by models themselves. Plus routing loses context making it more expensive and less reliable. Model Routing is just Bitter lesson. The models themselves will get better at this and frontier companies will simply give that capability

> Model Routing will always be done better by models themselves. > The models themselves will get better at this and frontier companies will simply give that capability I would never trust something like model routing to the same company that would profit from it, and that goes for telling models doing their own routing when that could easily be trained into the model to make things more expensive. Sort of a conflict…

Any different company providing the routing can be financially or otherwise incentivized to route this way or another.

Re: Claude Opus 5

#967

Earlier quoted context omitted.

It's likely bumping up against people's desire to have the model complete a given task without asking the person to intervene a bunch of times. Seems unclear how you satisfy everyone here.

Give the model judgement

Whose judgement?

I'm not a fan of how many "I set budget X and woke up to an eleventy trillion dollar bill" posts I see, and those are all generated by companies giving their tools the 'judgement' that the completion of the task is more important than the users wallet (or more cynically that they think they can actually get paid by just blasting unintended compute)

Re: Claude Opus 5

#969
post #845

"Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part." How surreal is it that we are not absolutely jaw-dropped by these types o…

Isn't this jaw-droppingly the wrong answer?! If the model isn't given any way to directly view the image shouldn't it just reply with one sentence asking for permission. This reads as a model hyper-trained to burn tokens. Like suppose you issue a command to Opus or Fable which doesn't make sense and requires a lot of work. It will almost certainly not push back on your silly request and go ahead and burn as many toke…

I don't think it is. It's really easy to get a persistent, clever, hacky model to dial that down a bit and just come back to chat before stomping off into the woods.

If a model couldn't ever do that in the first place, it'll just get stuck.

I work in the "ZeroOne" space, working on concepts and prototypes for things that don't exist in market yet. Sometimes these models crank hard and immolate tokens while grounding themselves on expensive-to-ingest self-developed frameworks. If the results are well judged and the crank-turn latency is low, I'm okay with the cost as long as the model isn't wasting my time.

But when I want to do more boilerplate work, I turn down the model and thinking level and get more traditional about restraining action. For the really hard stuff, I reach for the models that will start a token bonfire in the back yard.

Re: Claude Opus 5

#970

My excitement about Anthropic had fabled-out dramatically when they suspended my pro account about two weeks ago within just 12 hours of fair use. I was really mind-blown when I tried Fable 5 for the first time to help me improve a game I was working on but shortly, they decided that I had a suspicious activity and suspended my account without a clear reason. I submitted a an appeal describing that I am 100% sure I h…

Are you outside the US? And/or are you using VPN? Those are the two things that come to my mind that can cause overzealous security monitoring to flag someone. Another explanation could be the content itself. Does your game have anything at all to do with computer hacking or sexual content? Is there graphic violent language? People commonly report being unable to use AI to work on such things due to guardrails.

Yes, I am outside the US and I don't use a VPN.

For the game, I believe there is nothing suspicious, at least to me. It's a small puzzle platformer game with basic mechanics and very simple graphics, no harm, no violence or inappropriate content. In fact, my first two playterers right now are my two sons!

Post reply on HN