Live data from Hacker News

Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

anthropic.com

61–70 of 758 posts

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#61
post #35
post #9

I still feel like the difference between Sonnet and Opus is a bit unclear. Somewhere on Anthropic's website it says that Opus is the most advanced, but on other parts it says Sonnet is the most advanced and also the fastest. The UI doesn't make the distinction clear either. Then on Perplexity, Perplexity says that Opus is the most advanced, compared to Sonnet. And finally, in the table in the blogpost, Opus isn't eve…

Opus hasn't yet gotten an update from 3 to 3.5, and if you line up the benchmarks, the Sonnet "3.5 New" model seems to beat it everywhere. I think they originally announced that Opus would get a 3.5 update, but with every product update they are doing I'm doubting it more and more. It seems like their strategy is to beat the competition on a smaller model that they can train/tune more nimbly and pair it with outside-…

> Opus hasn't yet gotten an update from 3 to 3.5, and if you line up the benchmarks, the Sonnet "3.5 New" model seems to beat it everywhere

Why isn't Anthropic clearer about Sonnet being better then? Why isn't it included in the benchmark if new Sonnet beats Opus? Why are they so ambiguous with their language?

For example, https://www.anthropic.com/api says:

> Sonnet - Our best combination of performance and speed for efficient, high-throughput tasks.

> Opus - Our highest-performing model, which can handle complex analysis, longer tasks with many steps, and higher-order math and coding tasks.

And Opus is above/after Sonnet. That to me implies that Opus is indeed better than Sonnet.

But then you go to https://docs.anthropic.com/en/docs/about-claude/models and it says:

> Claude 3.5 Sonnet - Most intelligent model

- Claude 3 Opus - Powerful model for highly complex tasks

Does that mean Sonnet 3.5 is better than Opus for even highly complex tasks, since it's the "most intelligent model"? Or just for everything except "highly complex tasks"

I don't understand why this seems purposefully ambiguous?

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#62
post #35
post #9

I still feel like the difference between Sonnet and Opus is a bit unclear. Somewhere on Anthropic's website it says that Opus is the most advanced, but on other parts it says Sonnet is the most advanced and also the fastest. The UI doesn't make the distinction clear either. Then on Perplexity, Perplexity says that Opus is the most advanced, compared to Sonnet. And finally, in the table in the blogpost, Opus isn't eve…

Opus hasn't yet gotten an update from 3 to 3.5, and if you line up the benchmarks, the Sonnet "3.5 New" model seems to beat it everywhere. I think they originally announced that Opus would get a 3.5 update, but with every product update they are doing I'm doubting it more and more. It seems like their strategy is to beat the competition on a smaller model that they can train/tune more nimbly and pair it with outside-…

Opus 3.5 will likely be the answer to GPT-5. Same with Gemini 1.5 Ultra.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#63
post #59

Earlier quoted context omitted.

It's just bizarre to force a computer to go through a GUI to use another computer. Of course it's going to be expensive.

With UIPath, Appian, etc. the whole field of RPA (robotic process automation) is a $XX billion industry that is built on that exact premise (that it's more feasible to do automation via GUIs than badly built/non-existing APIs). Depending on how many GUI actions correspond to one equivalent AI orchestrated API call, this might also not be too bad in terms of efficiency.

Most of the GUIs are Web pages, though, so you could just interact directly with an HTTP server and not actually render the screen.

Or you could teach it to hack into the backend and add an API...

Oh, and on edit, "bizarre" and "multi-billion-dollar-industry" are well known not to be mutually exclusive.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#64

Is there an easy way to use Claude as a Co-Pilot in VS Code? If it is better at coding, it would be great to have it integrated.

Cursor uses Claude as its base model.

There may be extensions for VScode to do it but it will never be allowed in Copilot unless MS and OpenAI have a falling out.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#65
post #41

Great progress from Anthropic! They really shouldn't change models from under the hood, however. A name should refer to a specific set of model weights, more or less. On the other hand, as long as its actually advancing the Pareto frontier of capability, re-using the same name means everyone gets an upgrade with no switching costs. Though, all said, Claude still seems to be somewhat of an insider secret. "ChatGPT" ha…

There was a recent article[0] trending on HN a about their revenue numbers, split by B2C vs B2B.

Based on it, it seems like Anthropic is 60% of OpenAI API-revenue wise, but just 4% B2C-revenue wise. Though I expect this is partly because the Claude web UI makes 3.5 available for free, and there's not that much reason to upgrade if you're not using it frequently.

[0]: https://www.tanayj.com/p/openai-and-anthropic-revenue-breakd...

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#66

Why not rev the numbers? "3.5" vs. "3.5 New" feels weird -- is there a particular reason why Anthropic doesn't want to call this 3.6 (or even 3.5.1)?

exactly my thought too, go up with the version number! Some negative examples: Claude Sonnet 3.5 for Workstations, Claude Sonnet 3.5 XP, Claude Sonnet 3.5 Max Pro, Claude Sonnet 3.5 Elite, Claude Sonnet 3.5 Ultra

[deleted]

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#67

Is there an easy way to use Claude as a Co-Pilot in VS Code? If it is better at coding, it would be great to have it integrated.

You can use it in Cursor - called "Cursor Tab"

IMO Cursor Tab performs much better than Co-Pilot, easily works through things that would cause Co-Pilot to get stuck, you should give it a try

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#68
post #41

Great progress from Anthropic! They really shouldn't change models from under the hood, however. A name should refer to a specific set of model weights, more or less. On the other hand, as long as its actually advancing the Pareto frontier of capability, re-using the same name means everyone gets an upgrade with no switching costs. Though, all said, Claude still seems to be somewhat of an insider secret. "ChatGPT" ha…

Traveling to the US recently, I was surprised to see Claude ads around the city/in the airport. It seems like they're investing on marketing there.

In my country I've never seen anyone mention them at all.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#69
post #40
post #28

Earlier quoted context omitted.

> Opus has been stuck on 3.0, so Sonnet 3.5 is better So for example, Perplexity is wrong here implying that Opus is better than Sonnet? https://i.imgur.com/N58I4PC.png

I think as of this announcement that is indeed outdated information.

So Opus that costs $15.00/$75.00 for 1mil tokens (input/output) is now worse than the model that costs $3.00/$15.00?

That's according to https://docs.anthropic.com/en/docs/about-claude/models which has "claude-3-5-sonnet-20241022" as the latest model (today's date)

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#70

Why not rev the numbers? "3.5" vs. "3.5 New" feels weird -- is there a particular reason why Anthropic doesn't want to call this 3.6 (or even 3.5.1)?

My guess is they didn't actually change the model, that's what the version number no change is conveying. They did some engineering around it to make it respond better, perhaps more resources or different prompts. Same cutoff date too.
Post reply on HN