Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

481–490 of 1001 posts

Re: Claude Sonnet 4.6

#481

It seems that extra-usage is required to use the 1M context window for Sonnet 4.6. This differs from Sonnet 4.5, which allows usage of the 1M context window with a Max plan. ``` /model claude-sonnet-4-6[1m] ⎿ API error: 429 {"type":"error","error": {"type":"rate_limit_error","message":"Extra usage is required for long context requests."},"request_id":"[redacted]"} ```

think that just needs extra usage enabled? or actually using extra usage?

i cant believe that havent updated their code yet to be able to handle the 1M context on subscription auth

Re: Claude Sonnet 4.6

#482
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

It's very simple: prompt injection is a completely unsolved problem. As things currently stand, the only fix is to avoid the lethal trifecta. Unfortunately, people really, really want to do things involving the lethal trifecta. They want to be able to give a bot control over a computer with the ability to read and send emails on their behalf. They want it to be able to browse the web for research while helping you wr…

even if you limit to 2/3 I think any sort of persistence that can be picked up by agents with the other 1 can lead to compromise, like a stored XSS.

Re: Claude Sonnet 4.6

#483
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Isn't "computer use" just interaction with a shell-like environment, which is routine for current agents?

This is being downvoted but it shouldn't be.

If the ultimate goal is having a LLM control a computer, round-tripping through a UX designed for bipedal bags of meat with weird jelly-filled optical sensors is wildly inefficient.

Just stay in the computer! You're already there! Vision-driven computer use is a dead end.

Re: Claude Sonnet 4.6

#484

Earlier quoted context omitted.

Does it matter? Really? I can type awful stuff into a word processor. That's my fault, not the programs. So if I can trick an LLM into saying awful stuff, whose fault is that? It is also just a tool...

What is the tool supposed to be used for? If I sell you a marvelous new construction material, and you build your home out of it, you have certain expectations. If a passer-by throws an egg at your house, and that causes the front door to unlock, you have reason to complain. I'm aware this metaphor is stupid. In this case, it's the advertised use cases. For the word processor we all basically agree on the boundaries…

Isn't it up to the user how they want to use the tool? Why are people so hell bent on telling others how to press their buttons in a word processor ( or anywhere else for that matter ). The only thing that it does, is raising a new batch of Florida men further detached from reality and consequences.

Re: Claude Sonnet 4.6

#485

How do people keep track of all these versions and releases of all these models and their pros/cons? Seems like a fulltime hobby to me. I'd rather just improve my own skills with all that time and energy

on a subscription you cant access all that many different options, so you just stay with whatever the newest is unless it doesnt work.

Re: Claude Sonnet 4.6

#486
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products" that companies produce is already just extra. They can get rid of 1/3-2/3s of their labor and make the same amount of money, why wouldn't they.

ZeroHedge on twitter said the following:

"According to the market, AI will disrupt everything... except labor, which magically will be just fine after millions are laid off."

Its also worth noting that if you can create a business with an LLM, so can everyone else. And sadly everyone has the same ideas, everyone ends up working on the same things causing competition to push margins to nothing. There's nothing special about building with LLMs as anyone can just copy you that has access to the same models and basic thought processes.

This is basic economics. If everyone had an oil well on their property that was affordable to operate the price of oil would be more akin to the price of water.

EDIT: Since people are focusing on my water analogy I mean:

If everyone has easy access to the same powerful LLMs that would just drive down the value you can contribute to the economy to next to nothing. For this reason I don't even think powerful and efficient open source models, which is usually the next counter argument people make, are necessarily a good thing. It strips people of the opportunity for social mobility through meritocratic systems. Just like how your water well isn't going to make your rich or allow you to climb a social ladder, because everyone already has water.

Re: Claude Sonnet 4.6

#487

Still fails the car wash question, I took the prompt from the title of this thread: https://news.ycombinator.com/item?id=47031580 The answer was "Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — you'd barely move before arriving. Walking takes less than a minute, and you can simply drive it through the wash and walk back home afterward." I've tried several other variant…

My human partner also failed the car wash question. I guess they didn’t put a lot of thoughts into it.

Re: Claude Sonnet 4.6

#488

Earlier quoted context omitted.

Anthropic was the first to spam reddit with fake users and posts, flooding and controlling their subreddit to be a giant sycophant. They nuked the internet by themselves. Basically they are the willing and happy instigators of the dead internet as long as they profit from it. They are by no means ethical, they are a for-profit company.

I actually agree with you, but I have no idea how one can compete in this playing field. The second there are a couple of bad actors in spammarketing, your hands are tied. You really can’t win without playing dirty. I really hate this, not justifying their behaviour, but have no clue how one can do without the other.

Its just law of the jungle all over again. Might makes right. Outcomes over means.

Game theory wise there is no solution except to declare (and enforce) spaces where leeching / degrading the environment is punished, and sharing, building, and giving back to the environment is rewarded.

Not financially, because it doesn't work that way, usually through social cred or mutual values.

But yeah the internet can no longer be that space where people mutually agree to be nice to each other. Rather utility extraction dominates—influencers, hype traders, social thought manipulators-and the rest of the world quietly leaves if they know what's good for them.

Lovely times, eh?

Re: Claude Sonnet 4.6

#489

Earlier quoted context omitted.

Isn't "computer use" just interaction with a shell-like environment, which is routine for current agents?

No. Computer use (to anthropic, as in the article) is an LLM controlling a computer via a video feed of the display, and controlling it with the mouse and keyboard.

> controlling a computer via a video feed of the display, and controlling it with the mouse and keyboard.

I guess that's one way to get around robots.txt. Claim that you would respect it but since the bot is not technically a crawler it doesn't apply. It's also an easier sell to not identify the bot in the user agent string because, hey, it's not a script, it's using the computer like a human would!

Re: Claude Sonnet 4.6

#490
post #477

Earlier quoted context omitted.

Consequence is they are now facing an issue of “cancer villages” where the soil and water are unbelievably poisonous in many places.

which isnt particularly unique. its comparable to something like aome subset of americans getting black lung, or the health problems from the train explosion in east palestine. it took a lot of work for environmentalists to get some regulation into the US, canda, and the EU. china will get to that eventually

It isn’t. I just bring it up to state there is a very good reason the rest of the world doesn’t just drop their regulations. In the future I imagine China may give up many of these industries and move to cleaner ones, letting someone else take the toxic manufacturing.
Post reply on HN