Live data from Hacker News

Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

anthropic.com

181–190 of 758 posts

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#181
post #80

Seems like both: - AI Labs will eat some of the wrappers on top of their APIs - even complex ones like this. There are whole startups that are trying to build computer use. - AI is fitting _some_ scaling law - the best models are getting better and the "previously-state-of-the-art" models are fractions of what they cost a couple years ago. Though it remains to be seen if it's like Moore's Law or if incremental improv…

It seems a little silly to pretend there’s a scaling “law” without plotting any points or doing a projection. Without the mathiness, we could instead say that new models keep getting better and we don’t know how long that trend will continue.

> It seems a little silly to pretend there’s a scaling “law” without plotting any points or doing a projection.

Isn't this Kaplan 2020 or Hoffmann 2022?

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#182
One of the funnier things during training with the new API (which can control your computer) was this:

"Even while recording these demos, we encountered some amusing moments. In one, Claude accidentally stopped a long-running screen recording, causing all footage to be lost.

Later, Claude took a break from our coding demo and began to peruse photos of Yellowstone National Park."

[0] https://x.com/AnthropicAI/status/1848742761278611504

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#183

Earlier quoted context omitted.

It's just bizarre to force a computer to go through a GUI to use another computer. Of course it's going to be expensive.

Maybe fixing this for AI will finally force good accessibility support on major platforms/frameworks/apps (we can dream).

I really hope so. Even macOS voice control which has gotten pretty good is buggy with Messages, which is a core Apple app.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#184
post #10

How does the computer use work -- Is this a desktop app they are providing that can do actions on your computer? Didn't see any such mention in the post

It’s a sandbox compute environment, using Gvisor or Firecracker or similar, which exposes a browser environment to the LLM. modal.com’s modal.Sandbox can be the compute layer for this. It uses Gvisor under the hood.

Is there any Python/Node.js library to easily spawn secure isolated compute environments, possibly using gvisor or firecracker under the hood?

This could be useful to build a self-hosted "Computer use" using Ollama and a multimodal model.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#185
post #35

Earlier quoted context omitted.

Opus hasn't yet gotten an update from 3 to 3.5, and if you line up the benchmarks, the Sonnet "3.5 New" model seems to beat it everywhere. I think they originally announced that Opus would get a 3.5 update, but with every product update they are doing I'm doubting it more and more. It seems like their strategy is to beat the competition on a smaller model that they can train/tune more nimbly and pair it with outside-…

Opus 3.5 will likely be the answer to GPT-5. Same with Gemini 1.5 Ultra.

Maybe - would make sense not to release their latest greatest (Opus 4.0) until competition forces them to, and Amodei has previously indicated that they would rather respond to match frontier SOTA than themselves accelerate the pace of advance by releasing first.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#186

I have been a paying ChatGPT customer for a long time (since the very beginning). Last week I've compared ChatGPT to Claude and the results (to my eye) were better, the output better structured and the canvas works better. I'm on the edge of jumping ship.

interesting. i couldn’t imagine giving up o1-preview right now even with just 30/week.

and i do get a some bit of value from advanced voice mode, although it would be a lot more if it were unlimited

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#187
post #61
post #35

Earlier quoted context omitted.

Opus hasn't yet gotten an update from 3 to 3.5, and if you line up the benchmarks, the Sonnet "3.5 New" model seems to beat it everywhere. I think they originally announced that Opus would get a 3.5 update, but with every product update they are doing I'm doubting it more and more. It seems like their strategy is to beat the competition on a smaller model that they can train/tune more nimbly and pair it with outside-…

> Opus hasn't yet gotten an update from 3 to 3.5, and if you line up the benchmarks, the Sonnet "3.5 New" model seems to beat it everywhere Why isn't Anthropic clearer about Sonnet being better then? Why isn't it included in the benchmark if new Sonnet beats Opus? Why are they so ambiguous with their language? For example, https://www.anthropic.com/api says: > Sonnet - Our best combination of performance and speed fo…

> Why isn't Anthropic clearer about Sonnet being better then?

They are clear that both: Opus > Sonnet and 3.5 > 3.0. I don't think there is a clear universal better/worse relationship between Sonnet 3.5 and Opus 3.0; which is better is task dependent (though with Opus 3.0 being five times as expensive as Sonnet 3.5, I wouldn't be using Opus 3.0 unless Sonnet 3.5 proved clearly inadequate for a task.)

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#188
post #182

One of the funnier things during training with the new API (which can control your computer) was this: "Even while recording these demos, we encountered some amusing moments. In one, Claude accidentally stopped a long-running screen recording, causing all footage to be lost. Later, Claude took a break from our coding demo and began to peruse photos of Yellowstone National Park." [0] https://x.com/AnthropicAI/status/1…

At least now we know SkyClaude’s plan to end human civilization.

It’s planning on triggering a Yellowstone caldera super eruption.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#189
post #85

Earlier quoted context omitted.

> I don't understand why this seems purposefully ambiguous? I wouldn't attribute this to malice when it can also be explained by incompetence. Sonnet 3.5 New > Opus 3 > Sonnet 3.5 is generally how they stack up against each other when looking at the total benchmarks. "Sonnet 3.5 New" has just been announced, and they likely just haven't updated the marketing copy across the whole page yet, and maybe also haven't figu…

> B) potentially phase out Opus, and instead introduce new branding for what they called a "reasoning model" like OpenAI did with o1(-preview) When should we be using the -o OpenAI models? I've not been keeping up and the official information now assumes far too much familiarity to be of much use.

I think it's first important to note that there is a huge difference between -o models (GPT 4o; GPT 4o mini) and the o1 models (o1-preview; o1-mini).

The -o models are "just" stronger versions of their non-suffixed predecessors. They are the latest (and maybe last?) version of models in the lineage of GPT models (roughly GPT-1 -> GPT-2 -> GPT-3 -> GPT-3.5 -> GPT-4 -> GPT-4o).

The o1 models (not sure what the naming structure for upcoming models will be) are a new family of models that try to excel at deep reasoning, by allowing the models to use an internal (opaque) chain-of-thought to produce better results at the expense of higher token usage (and thus cost) and longer latency.

Personally, I think the use cases that justify the current cost and slowness of o1 are incredibly narrow (e.g. offline analysis of financial documents or deep academic paper research). I think in most interactive use-cases I'd rather opt for GPT-4o or Sonnet 3.5 instead of o1-preview and have the faster response time and send a follow-up message. Similarly for non-interactive use-cases I'd try to add a layer of tool calling with those faster models than use o1-preview.

I think the o1-like models will only really take off, if the prices for it are coming down, and it is clearly demonstrated that more "thinking tokens" correlate to predictably better results, and results that can compete with highly tuned prompts/fine tuned models that or currently expensive to produce in terms of development time.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#190
post #134

Earlier quoted context omitted.

I have to agree. I've been chatting with Claude for the first time in a couple days and while it's very on-par with ChatGPT 4o in terms of capability, it has this difficult-to-quantify feeling of being warmer and friendlier to interact with. I think the human name, serif font, system prompt, and tendency to create visuals contributes to this feeling.

>it's very on-par with ChatGPT 4o in terms of capability The previous 3.5 Sonnet checkpoint was already better than GPT-4o in terms of programming and multi-language capabilities. Also, GPT-4o sometimes feels completely moronic, for example, the other day I asked for fun a technical question about configuring a "dream-sync" device to comply with the "Personal Consciousness Data Protection Act", and GPT-4o just replie…

actually, that's what makes chat gpt powerful. I like an LLM willing to go along with what ever I am trying to do, because one day I might be coding, and another day I might be just trying to role play, write a book, what ever.

I really cant understand what you were expecting, a tool works with how you use it, if you smack a hammer into your face, don't complain about a bloody nose. maybe dont do like that?

Post reply on HN