Earlier quoted context omitted.
Until 2 remain, then it's extraction time.
Or self host the oss models on the second hand GPU and RAM that's left when the big labs implode
Claude Sonnet 4.6
421–430 of 1001 posts
Re: Claude Sonnet 4.6
#422Earlier quoted context omitted.
Did you see the graph benchmark? I found it quite interesting. It had to do a graph traversal on a natural text representation of a graph. Pretty much your problem.
Update: I took a corpus of personal chat data (this way it wouldn't be seen in training), and tried asking it some paraphrased questions. It performed quite poorly.
Re: Claude Sonnet 4.6
#423Earlier quoted context omitted.
Remember when GPT-2 was “too dangerous to release” in 2019? That could have still been the state in 2026 if they didn’t YOLO it and ship ChatGPT to kick off this whole race.
I was just thinking earlier today how in an alternate universe, probably not too far removed from our own, Google has a monopoly on transformers and we are all stuck with a single GPT-3.5 level model, and Google has a GPT-4o model behind the scenes that it is terrified to release (but using heavily internally).
Where would we be if patents never existed?
Re: Claude Sonnet 4.6
#424Earlier quoted context omitted.
“Cryptic” exit posts are basically noise. If we are going to evaluate vendors, it should be on observable behavior and track record: model capability on your workloads, reliability, security posture, pricing, and support. Any major lab will have employees with strong opinions on the way out. That is not evidence by itself.
We recently had an employee leave our team, posting an extensive essay on LinkedIn, "exposing" the company and claiming a whole host of wrong-doing that went somewhat viral. The reality is, she just wasn't very good at her job and was fired after failing to improve following a performance plan by management. We all knew she was slacking and despite liking her on a personal level, knew that she wasn't right for what i…
Re: Claude Sonnet 4.6
#425I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…
Isn't "computer use" just interaction with a shell-like environment, which is routine for current agents?
Computer use (to anthropic, as in the article) is an LLM controlling a computer via a video feed of the display, and controlling it with the mouse and keyboard.
Re: Claude Sonnet 4.6
#426Earlier quoted context omitted.
An Anthropic safety researcher just recently quit with very cryptic messages , saying "the world is in peril"... [1] (which may mean something, or nothing at all) Codex quite often refuses to do "unsafe/unethical" things that Anthropic models will happily do without question. Anthropic just raised 30 bn... OpenAI wants to raise 100bn+. Thinking any of them will actually be restrained by ethics is foolish. [1] https:/…
I think we're fine: https://youtube.com/shorts/3fYiLXVfPa4?si=0y3cgdMHO2L5FgXW Claude invented something completely nonsensical: > This is a classic upside-down cup trick! The cup is designed to be flipped — you drink from it by turning it upside down, which makes the sealed end the bottom and the open end the top. Once flipped, it functions just like a normal cup. *The sealed "top" prevents it from spilling while it…
Re: Claude Sonnet 4.6
#427I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…
Isn't "computer use" just interaction with a shell-like environment, which is routine for current agents?
> where the model interacts with the GUI (graphical userinterface) directly.
Re: Claude Sonnet 4.6
#428I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…
Isn't "computer use" just interaction with a shell-like environment, which is routine for current agents?
> hundreds of tasks across real software (Chrome, LibreOffice, VS Code, and more) running on a simulated computer. There are no special APIs or purpose-built connectors; the model sees the computer and interacts with it in much the same way a person would: clicking a (virtual) mouse and typing on a (virtual) keyboard.
Re: Claude Sonnet 4.6
#429I'm a bit surprised it gets this question wrong (ChatGPT gets it right, even on instant). All the pre-reasoning models failed this question, but it's seemed solved since o1, and Sonnet 4.5 got it right. https://claude.ai/share/876e160a-7483-4788-8112-0bb4490192af This was sonnet 4.6 with extended thinking.