Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

571–580 of 1001 posts

Re: Claude Sonnet 4.6

#571
post #469

Earlier quoted context omitted.

Remarkable, since the goal is clearly stated and the language isn’t tricky.

Well it is a trick question due to it being non-sensical. The AI is interpreting it in the only way that makes sense, the car is already at the car wash, should you take a 2nd car to the car wash 50 meters away or walk. It should just respond "this question doesn't make any sense, can you rephrase it or add additional information"

The question isn't nonsense, it just has an answer which is so obvious nobody would ever ask it organically.

Re: Claude Sonnet 4.6

#572
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

It does not seem all that problematic for the most obviously valuable use case: You use an (web) app, that you consider reasonably safe, but that offers no API, and you want to do things with it. The whole adversarial action problem just dissipates, because there is no adversary anywhere in the path.

No random web browsing. Just opening the same app, every day. Login. Read from a calendar or a list. Click a button somewhere when x == true. Super boring stuff. This is an entire class of work that a lot of humans do in a lot of companies today, and there it could be really useful.

Re: Claude Sonnet 4.6

#573
post #559

Still fails the car wash question, I took the prompt from the title of this thread: https://news.ycombinator.com/item?id=47031580 The answer was "Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — you'd barely move before arriving. Walking takes less than a minute, and you can simply drive it through the wash and walk back home afterward." I've tried several other variant…

Looking at the responses below it's interesting how binary they are. It's classic hallucinations style where it's flopping between two alternatives but which ever one it picks it's absolutely confident about.

You can always make it go back and forth with "Are you sure?".

The fact that these are still issues ~6 years into this tech is bewildering.

Re: Claude Sonnet 4.6

#574

It excels at agentic knowledge work. These custom, domain-specific playbooks are tailor made: claudecodehq.com

Is this technique of spamming with vibe-coded “directories” really working? Genuinely curious

We have to start banning users who do this

Re: Claude Sonnet 4.6

#575

Earlier quoted context omitted.

What is the tool supposed to be used for? If I sell you a marvelous new construction material, and you build your home out of it, you have certain expectations. If a passer-by throws an egg at your house, and that causes the front door to unlock, you have reason to complain. I'm aware this metaphor is stupid. In this case, it's the advertised use cases. For the word processor we all basically agree on the boundaries…

Isn't it up to the user how they want to use the tool? Why are people so hell bent on telling others how to press their buttons in a word processor ( or anywhere else for that matter ). The only thing that it does, is raising a new batch of Florida men further detached from reality and consequences.

Users can use tools how they want. However, some of those uses are hazards. If I am trying to scare birds away from my house with fireworks and burn my neighbors' house down, that's kind of a problem for me. If these fireworks are marketed as practical bird repellent, that's a problem for me and the manufacturer.

I'm not sure if it's official marketing or just breathless hype men or an astroturf campaign.

Re: Claude Sonnet 4.6

#576
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

It does not seem all that problematic for the most obviously valuable use case: You use an (web) app, that you consider reasonably safe, but that offers no API, and you want to do things with it. The whole adversarial action problem just dissipates, because there is no adversary anywhere in the path. No random web browsing. Just opening the same app, every day. Login. Read from a calendar or a list. Click a button so…

You're maybe used to a world in which we've gotten rid of in-band signaling and XSS and such, so if I write you a check and put the string "Memo'); DROP TABLE accounts; --" [0] or "" in the memo, you might see that text on your bank's website.

But LLM's are back to the old days of in-band signaling. If you have an LLM poking at your bank's website for you, and I write you a check with a memo containing the prompt injection attack du jour, your LLM will read it. And the whole point of all these fancy agentic things is that they're supposed to have the freedom to do what they think is useful based on the information available to them. So they might follow the directions in the memo field.

Or the instructions in a photo on a website. Or instructions in an ad. Or instructions in an email. Or instructions in the Zelle name field for some other user. Or instructions in a forum post.

You show me a website where 100% of the content, including the parts that are clearly marked (as a human reader) as being from some other party, is trustworthy, and I'll show you a very boring website.

(Okay, I'm clearly lying -- xkcd.org is open and it's pretty much a bunch of static pages that only have LLM-readable instructions in places where the author thought it would be funny. And I guess if I have an LLM start poking at xkcd.org for me, I deserve whatever happens to me. I have one other tab open that probably fits into this probably-hard-to-prompt-inject open, and it is indeed boring and I can't think of any reason that I would give an LLM agent with any privileges at all access to it.)

[0] https://xkcd.com/327/

Re: Claude Sonnet 4.6

#577
post #231

Earlier quoted context omitted.

For comparisonI think the current leader in pelican drawing is Gemini 3 Deep Think: https://bsky.app/profile/simonwillison.net/post/3meolxx5s722...

My take (also Gemini 3 Deep Think): https://gemini.google.com/share/12e672dd39b7 Somehow it's much better now.

Is that actually better? That pelican has arms sprouting out of its wings

Re: Claude Sonnet 4.6

#578

Earlier quoted context omitted.

OK! customisations please? ...

All of my “characteristics” (a setting I don’t think I’ve seen before) are set to default and my custom instructions are as follows… —— Always assume British English when relevant. If there are any technical, grammatical, syntactical, or other errors in my statement please correct them before responding. Tell it like it is; don't sugar-coat responses. Adopt a skeptical, questioning approach.

Hah, your experience is a great example of the futility of recommendations to add instructions to "solve" issues like sycophancy, just trading one form of insufferable chatbot for something even more insufferable. Different strokes and all but there's no way I could tolerate reading that every day, particularly when it's completely wrong...

Re: Claude Sonnet 4.6

#579

Earlier quoted context omitted.

Does it matter? How much power does it take to run duolingo? How much power did it take to manufacture 300000 Teslas? Everything takes power

I think it does matter how much power it takes but, in the context of power to "benefits humanity" ratio. Things that significantly reduce human suffering or improve human life are probably worth exerting energy on. However, if we frame the question this way, I would imagine there are many more low-hanging fruit before we question the utility of LLMs. For example, should some humans be dumping 5-10 kWh/day into thing…

[dead]

Re: Claude Sonnet 4.6

#580
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

This is the elephant in the room nobody wants to talk about. AI is dead in the water for the supposed mass labor replacement that will happen unless this is fixed. Summarize some text while I supervise the AI = fine and a useful productivity improvement, but doesn’t replace my job. Replace me with an AI to make autonomous decisions outside in the wild and liability-ridden chaos ensues. No company in their right mind…

And why would it materialize? Anyone who has used even modern models like Opus 4.6 in very long and extensive chats about concrete topics KNOWS that this LLM form of Artificial Intelligence is anything but intelligent.

You can see the cracks happening quite fast actually and you can almost feel how trained patterns are regurgitated with some variance - without actually contextualizing and connecting things. More guardrailing like web sources or attachments just narrow down possible patterns but you never get the feeling that the bot understands. Your own prompting can also significantly affect opinions and outcomes no matter the factual reality.

Post reply on HN