Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

421–430 of 1001 posts

Re: Claude Sonnet 4.6

#422

Earlier quoted context omitted.

Did you see the graph benchmark? I found it quite interesting. It had to do a graph traversal on a natural text representation of a graph. Pretty much your problem.

Update: I took a corpus of personal chat data (this way it wouldn't be seen in training), and tried asking it some paraphrased questions. It performed quite poorly.

Which models did you try?

Re: Claude Sonnet 4.6

#423

Earlier quoted context omitted.

Remember when GPT-2 was “too dangerous to release” in 2019? That could have still been the state in 2026 if they didn’t YOLO it and ship ChatGPT to kick off this whole race.

I was just thinking earlier today how in an alternate universe, probably not too far removed from our own, Google has a monopoly on transformers and we are all stuck with a single GPT-3.5 level model, and Google has a GPT-4o model behind the scenes that it is terrified to release (but using heavily internally).

Now think about how often the patent system has stifled and stalled and delayed advancement for decades per innovation at a time.

Where would we be if patents never existed?

Re: Claude Sonnet 4.6

#424

Earlier quoted context omitted.

“Cryptic” exit posts are basically noise. If we are going to evaluate vendors, it should be on observable behavior and track record: model capability on your workloads, reliability, security posture, pricing, and support. Any major lab will have employees with strong opinions on the way out. That is not evidence by itself.

We recently had an employee leave our team, posting an extensive essay on LinkedIn, "exposing" the company and claiming a whole host of wrong-doing that went somewhat viral. The reality is, she just wasn't very good at her job and was fired after failing to improve following a performance plan by management. We all knew she was slacking and despite liking her on a personal level, knew that she wasn't right for what i…

Same thing happened to us. Me and a C level guy were personally attacked. It feels really bad to see someone you actually tried really hard to help fit in , but just couldn’t despite really wanting the person to succeed, come around and accuse you of things that clearly aren’t true. HR got the to remove the “review” eventually but now there’s a little worry about what the team really thinks, whether they would do the same in some future layoff (we never had any, the person just wasn’t very good).

Re: Claude Sonnet 4.6

#425
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Isn't "computer use" just interaction with a shell-like environment, which is routine for current agents?

No.

Computer use (to anthropic, as in the article) is an LLM controlling a computer via a video feed of the display, and controlling it with the mouse and keyboard.

Re: Claude Sonnet 4.6

#426

Earlier quoted context omitted.

An Anthropic safety researcher just recently quit with very cryptic messages , saying "the world is in peril"... [1] (which may mean something, or nothing at all) Codex quite often refuses to do "unsafe/unethical" things that Anthropic models will happily do without question. Anthropic just raised 30 bn... OpenAI wants to raise 100bn+. Thinking any of them will actually be restrained by ethics is foolish. [1] https:/…

I think we're fine: https://youtube.com/shorts/3fYiLXVfPa4?si=0y3cgdMHO2L5FgXW Claude invented something completely nonsensical: > This is a classic upside-down cup trick! The cup is designed to be flipped — you drink from it by turning it upside down, which makes the sealed end the bottom and the open end the top. Once flipped, it functions just like a normal cup. *The sealed "top" prevents it from spilling while it…

He tried this with ChatGPT too. It called the item a "novelty cup" you couldn't drink out of :)

Re: Claude Sonnet 4.6

#427
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Isn't "computer use" just interaction with a shell-like environment, which is routine for current agents?

No their definition of "computer use" now means:

> where the model interacts with the GUI (graphical userinterface) directly.

Re: Claude Sonnet 4.6

#428
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Isn't "computer use" just interaction with a shell-like environment, which is routine for current agents?

> Almost every organization has software it can’t easily automate: specialized systems and tools built before modern interfaces like APIs existed. [...]

> hundreds of tasks across real software (Chrome, LibreOffice, VS Code, and more) running on a simulated computer. There are no special APIs or purpose-built connectors; the model sees the computer and interacts with it in much the same way a person would: clicking a (virtual) mouse and typing on a (virtual) keyboard.

https://www.anthropic.com/news/claude-sonnet-4-6

Re: Claude Sonnet 4.6

#429

I'm a bit surprised it gets this question wrong (ChatGPT gets it right, even on instant). All the pre-reasoning models failed this question, but it's seemed solved since o1, and Sonnet 4.5 got it right. https://claude.ai/share/876e160a-7483-4788-8112-0bb4490192af This was sonnet 4.6 with extended thinking.

[deleted]
Post reply on HN