Live data from Hacker News

Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

anthropic.com

101–110 of 758 posts

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#101

Is there an easy way to use Claude as a Co-Pilot in VS Code? If it is better at coding, it would be great to have it integrated.

Codeium (cheapest), double.bot and continue.dev (with api key) have Claude in chat.

https://github.com/cline/cline (with api key) has Claude as agent.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#102
post #85
post #61

Earlier quoted context omitted.

> Opus hasn't yet gotten an update from 3 to 3.5, and if you line up the benchmarks, the Sonnet "3.5 New" model seems to beat it everywhere Why isn't Anthropic clearer about Sonnet being better then? Why isn't it included in the benchmark if new Sonnet beats Opus? Why are they so ambiguous with their language? For example, https://www.anthropic.com/api says: > Sonnet - Our best combination of performance and speed fo…

> I don't understand why this seems purposefully ambiguous? I wouldn't attribute this to malice when it can also be explained by incompetence. Sonnet 3.5 New > Opus 3 > Sonnet 3.5 is generally how they stack up against each other when looking at the total benchmarks. "Sonnet 3.5 New" has just been announced, and they likely just haven't updated the marketing copy across the whole page yet, and maybe also haven't figu…

> I wouldn't attribute this to malice when it can also be explained by incompetence.

I don't think it's malice either, but if Opus costs more to them to run, and they've already set a price they cannot raise, it makes sense they want people to use models they have a higher net return on, that's just "business sense" and not really malice.

> and they likely just haven't updated the marketing copy across the whole page yet

The API docs have been updated though, which is the second page I linked. It mentions the new model by it's full name "claude-3-5-sonnet-20241022" so clearly they've gone through at least that page. Yet the wording remains ambiguous.

> Sonnet 3.5 New > Opus 3 > Sonnet 3.5 is generally how they stack up against each other when looking at the total benchmarks.

Which ones are you looking at? Since the benchmark comparison in the blogpost itself doesn't include Opus at all.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#103

Completely irrelevant, and it might just be me, but I really like Anthropic's understated branding. OpenAI's branding isn't exactly screaming in your face either, but for something that's generated as much public fear/scaremongering/outrage as LLMs have over the last couple of years, Anthropic's presentation has a much "cosier" veneer to my eyes. This isn't the Skynet Terminator wipe-us-all-out AI, it's the adorable…

Anthropic has recently begun a new, big ad campaign (ads in Times Square) that more-or-less takes potshots at OpenAI. https://www.reddit.com/r/singularity/comments/1g9e0za/anthro...

[deleted]

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#104
Claude's current ability to use computers is imperfect. Some actions that people perform effortlessly—scrolling, dragging, zooming—currently present challenges for Claude and we encourage developers to begin exploration with low-risk tasks.

Nice, but I wonder why didn't they use UI automation/accessibility libraries, that have access to the semantic structure of apps/web pages, as well as accessing documents directly instead of having Excel display them for you.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#106

Is there an easy way to use Claude as a Co-Pilot in VS Code? If it is better at coding, it would be great to have it integrated.

Tabnine includes Claude as an option. I've been using it to compare Claude Sonnet to Chatgpt-4o and Sonnet is clearly much better.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#107
post #69
post #40

Earlier quoted context omitted.

I think as of this announcement that is indeed outdated information.

So Opus that costs $15.00/$75.00 for 1mil tokens (input/output) is now worse than the model that costs $3.00/$15.00? That's according to https://docs.anthropic.com/en/docs/about-claude/models which has "claude-3-5-sonnet-20241022" as the latest model (today's date)

Yes, you will find similar things at essentially all other model providers.

The older/bigger GPT4 runs at $30/$60 and peforms about on par with GPT4o-mini which costs only $0.15/$0.60.

If you are currently, or have been integrating AI models in the past ~2 years, you should definitely keep up with model capability/pricing development. If you are staying on old models you are certainly overpaying/leaving performance on the table. It's essentially a tax on agility.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#108
post #80

Seems like both: - AI Labs will eat some of the wrappers on top of their APIs - even complex ones like this. There are whole startups that are trying to build computer use. - AI is fitting _some_ scaling law - the best models are getting better and the "previously-state-of-the-art" models are fractions of what they cost a couple years ago. Though it remains to be seen if it's like Moore's Law or if incremental improv…

It seems a little silly to pretend there’s a scaling “law” without plotting any points or doing a projection. Without the mathiness, we could instead say that new models keep getting better and we don’t know how long that trend will continue.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#109
post #92
post #57

Earlier quoted context omitted.

It worked for Nintendo. The 3ds and “new 3ds” were both big sellers.

3ds doesn't have a version number to bump. Claude 3.5 does.

I hear the Nintendo 4DS was very popular with the higher dimensional beings!

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#110

This is what the Rabbit "large action model" pretended to be. Wouldn't be surprised to see them switch to this and claim they were never lying about their capabilities because it works now. Pretty cool for sure.

I think Rabbit had the business model wrong though, I don't think automating UI's to order pizza is anywhere near as valuable as automating the app workflows for B2B users.
Post reply on HN