Live data from Hacker News

Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

anthropic.com

211–220 of 758 posts

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#211
post #99

my quick notes on Computer Use: - "computer use" is basically using Claude's vision + tool use capability in a loop. There's a reference impl but there's no "claude desktop" app that just comes with this OOTB - they're basically advertising that they bumped up Claude 3.5's screen vision capability. we discussed the importance of this general computer agent approach with David on our pod https://x.com/swyx/status/1771…

Haven't used vision models before, can someone comment if they are good at "pointing things". E.g given a picture, give co-ordinate for text "foo".

This is the key to accurate control, it needs to be very precise.

Maybe Claude's model is trained at this. Also what about open source vision models? Any ones good at "pointing things" on a typical computer screen?

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#212

Earlier quoted context omitted.

OpenAI does exactly the same thing, by the way; the named models also have dated versions. For instance, there current models include (only listing versions with more than one dated version for the same "name" version): gpt-4o-2024-08-06 gpt-4o-2024-05-13 gpt-4-0125-preview gpt-4-1106-preview gpt-4-0613 gpt-4-0314 gpt-3.5-turbo-0125 gpt-3.5-turbo-1106

On the one hand, if OpenAI makes a bad choice, it’s still a bad choice to copy it. On the other hand, OpenAI has moved to a naming convention where they seem to use a name for the model: “GPT-4”, “GPT-4 Turbo”, “GPT-4o”, “GPT-4o mini”. Separately, they use date strings to represent the specific release of that named model. Whereas Anthropic had a name: “Claude Sonnet”, and what appeared to be an incrementing version…

> Now, Anthropic is jamming two version strings on the same product, and I consider that a bad choice. It doesn’t mean I think OpenAI’s approach is great either, but I think there are nuances that say they’re not doing exactly the same thing

Anthropic has always had dated versions as well as the other components, and they are, in fact, doing exactly the same thing, except that OpenAI has a base model in each generation with no suffix before the date specifier (what I call the "Model Class" on the table below), and OpenAI is inconsistent in their date formats, see:

  Major Family  Generation    Model Class Date
  claude        3.5           sonnet      20041022
  claude        3.0           opus        20240229
  gpt           4             o           2024-08-06
  gpt           4             o-mini      2024-07-18
  gpt           4             -           0613
  gpt           3.5           turbo       0125

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#213

I've seen quite a few YC startups working on AI-powered RPA, and now it looks like a foundational model player is directly competing in their space. It will be interesting to see whether Anthropic will double down on this or leave it to third-party developers to build commercial applications around it.

We're one of those players (https://github.com/Skyvern-AI/skyvern) and we're definitely watching the space with a lot of excitement

We thought it was inevitable that OpenAI / Anthropic would veer into this space and start to become competitive with us. We actually expected OpenAI to do it first!

What this confirms is that there is significant interest in computer / browser automation, and the problem is still unsolved. We will see whether the automation itself is an application later problem (our approach) or whether the model needs to be intertwined with the application (Anthropic's approach here)

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#214

Earlier quoted context omitted.

On the one hand, if OpenAI makes a bad choice, it’s still a bad choice to copy it. On the other hand, OpenAI has moved to a naming convention where they seem to use a name for the model: “GPT-4”, “GPT-4 Turbo”, “GPT-4o”, “GPT-4o mini”. Separately, they use date strings to represent the specific release of that named model. Whereas Anthropic had a name: “Claude Sonnet”, and what appeared to be an incrementing version…

> Now, Anthropic is jamming two version strings on the same product, and I consider that a bad choice. It doesn’t mean I think OpenAI’s approach is great either, but I think there are nuances that say they’re not doing exactly the same thing Anthropic has always had dated versions as well as the other components, and they are, in fact, doing exactly the same thing, except that OpenAI has a base model in each generati…

But did they ever have more than one release of Claude 3 Sonnet? Or any other model prior to today?

As far as I can tell, the answer is “no”. If true, then the fact that they previously had date strings would be a purely academic footnote to what I was saying, not actually relevant or meaningful.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#215
anybody know how the hell they're combating / gonna combat captcha's, cloudflare blocking, etc. I remember playing in this space on a toy project and being utterly frustrated by anti-scraping. Maybe one good thing that will come out of this AI boom is that companies will become nicer to scrapers? Or maybe, they'll just cut sweetheart deals?

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#216
And today I realized that despite it being an extremely common activity, we don’t really have a word for “using the computer” which is distinct from “computing”. It’s funny because AI models are always “using a computer” but now they can “use your computer.”

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#217

Completely irrelevant, and it might just be me, but I really like Anthropic's understated branding. OpenAI's branding isn't exactly screaming in your face either, but for something that's generated as much public fear/scaremongering/outrage as LLMs have over the last couple of years, Anthropic's presentation has a much "cosier" veneer to my eyes. This isn't the Skynet Terminator wipe-us-all-out AI, it's the adorable…

Take a read through the user agreements for all the major LLM providers and marvel at the simplicity and customer friendliness of the Anthropic one vs the others.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#218
Both new Sonnet and gpt-4o still fail at a simple:

"How many w's are in strawberry?"

gpt-4o: There are 2 "w's" in "strawberry."

Claude 3.5 Sonnet (new): Let me count the w's in "strawberry": 0 w's.

(same question with 'r' succeeds)

What is artificial about current gen of "artificial intelligence" is the way training (predict next token) and benchmarking (overfitting) is done. Perhaps a fresh approach is needed to achieve a true next step.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#219
post #134

Earlier quoted context omitted.

>it's very on-par with ChatGPT 4o in terms of capability The previous 3.5 Sonnet checkpoint was already better than GPT-4o in terms of programming and multi-language capabilities. Also, GPT-4o sometimes feels completely moronic, for example, the other day I asked for fun a technical question about configuring a "dream-sync" device to comply with the "Personal Consciousness Data Protection Act", and GPT-4o just replie…

actually, that's what makes chat gpt powerful. I like an LLM willing to go along with what ever I am trying to do, because one day I might be coding, and another day I might be just trying to role play, write a book, what ever. I really cant understand what you were expecting, a tool works with how you use it, if you smack a hammer into your face, don't complain about a bloody nose. maybe dont do like that?

So if you're trying to write code and mistakenly ask it how to use a nonexistent API, you'd rather it give you garbage rather than explaining your mistake and helping you fix it? After all, you're clearly just roleplaying, right?
Post reply on HN