Live data from Hacker News

Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

anthropic.com

141–150 of 758 posts

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#142
post #15

Earlier quoted context omitted.

For a company selling intelligence, that's a pretty stupid way of labelling a new product.

"computer use" is also as bad a marketing choice as possible for something that actually seems pretty cool.

it makes sense in contrast to "tool use". basically, either fly-by-vision or fly-by-instruments, same dilemma you have in self driving cars

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#143
post #87

Of course there's great inefficiency in having the Claude software control a computer with a human GUI mediating everything, but it's necessary for many uses right now given how much we do where only human interfaces are easily accessible. If something like it takes off, I expect interfaces for AI software would be published, standardized, etc. Your customers may not buy software that lacks it. But what I really want…

I agree, I bet models could excel at CLI tasks since the feedback would be immediate and in a language they can readily consume. It's probably much easier for them to to handle "command requires 2 arguments and only 1 was provided" than to do image-to-text on an error modal and apply context to figure out what went wrong.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#144
post #89

Earlier quoted context omitted.

It's just bizarre to force a computer to go through a GUI to use another computer. Of course it's going to be expensive.

Building an entirely new world for agents to compute in is far more difficult than building an agent that can operate in a human world. However i'm sure over time people will start building bridges to make it easier/cheaper for agents to operate in their own native environment. It's like another digital transformation. Paper lasted for years before everything was digitalized. Human interfaces will last for years befo…

I am just a dilettante, but I imagined that eventually agents will be making API calls directly via browser extension, or headless browser.

I assumed everyone making these UI agents will create a library of each URL's API specification, trained by users.

Does that seem workable?

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#145

Completely irrelevant, and it might just be me, but I really like Anthropic's understated branding. OpenAI's branding isn't exactly screaming in your face either, but for something that's generated as much public fear/scaremongering/outrage as LLMs have over the last couple of years, Anthropic's presentation has a much "cosier" veneer to my eyes. This isn't the Skynet Terminator wipe-us-all-out AI, it's the adorable…

Anthropic has recently begun a new, big ad campaign (ads in Times Square) that more-or-less takes potshots at OpenAI. https://www.reddit.com/r/singularity/comments/1g9e0za/anthro...

Wonder what a normal person thinks this is an ad for

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#146

Claude is absurdly better at coding tasks than OpenAI. Like it's not even close. Particularly when it comes to hallucinations. Prompt for prompt, I see Claude being rock solid and returning fully executable code, with all the correct imports, while OpenAI struggles to even complete the task and will make up nonexistent libraries/APIs out of whole cloth.

Yeah, sonnet is noticeably better. To the point that openai is almost unusable, too many small errors

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#147
post #87

Of course there's great inefficiency in having the Claude software control a computer with a human GUI mediating everything, but it's necessary for many uses right now given how much we do where only human interfaces are easily accessible. If something like it takes off, I expect interfaces for AI software would be published, standardized, etc. Your customers may not buy software that lacks it. But what I really want…

I hope specialized interfaces for AI never happen. I want AI to use human interfaces, because I want to be empowered to use the same interfaces as AI in the future. A future where only AI can do things because it uses an incomprehensible special interface and the human interface is broken or non-existent is a dystopia. I also want humanoid robots instead of specialized non-humanoid robots for the same reason.

Maybe we'll end up with both, kind of like how we have scripting languages for ease of development, but we also can write assembly if we need bare metal access for speed.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#149
post #126

Earlier quoted context omitted.

Yes, you will find similar things at essentially all other model providers. The older/bigger GPT4 runs at $30/$60 and peforms about on par with GPT4o-mini which costs only $0.15/$0.60. If you are currently, or have been integrating AI models in the past ~2 years, you should definitely keep up with model capability/pricing development. If you are staying on old models you are certainly overpaying/leaving performance o…

> The older/bigger GPT4 runs at $30/$60 and peforms about on par with GPT4o-mini which costs only $0.15/$0.60. I don't think GPT-4o Mini has comparable performance to GPT-4 at all, where are you finding the benchmarks claiming this? Everywhere I look says GPT-4 is more powerful, but GPT-4o Mini is most cost-effective, if you're OK with worse performance. Even OpenAI themselves about GPT-4o Mini: > Our affordable and…

> Yeah, I mean that's why we're both here and why we're discussing this very topic, right? :D

That wasn't specifically directed at "you", but more as a plea to everyone reading that comment ;)

I looked at a few benchmarks, comparing the two, which like in the case of Opus 3 vs Sonnet 3.5 is hard, as the benchmarks the wider community is interested in shifts over time. I think this page[0] provides the best overview I can link to.

Yes, GPT4 is better in the MMLU benchmark, but in all other benchmarks and the LMSys Chatbot Arena scores[1], GPT4o-mini comes out ahead. Overall, the margin between is so thin that it falls under my definition of "on par". I think OpenAI is generally a bit more conservative with the messaging here (which is understandable), and they only advertise a model as "more capable", if one model beats the other one in every benchmark they track, which AFAIK is the case when it comes to 4o mini vs 3.5 Turbo.

[0]: https://context.ai/compare/gpt-4o-mini/gpt-4

[1]: https://artificialanalysis.ai/models?models_selected=gpt-4o-...

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#150

How long until "computer use" is tricked into entering PII or PHI into an attackers website?

I imagine initial computer use models will be kind of like untrained or unskilled computer users today (for example, some kids and grandparents). They'll do their best but will inevitably be easy to trick into clicking unscrupulous links and UI elements.

Will an AI model be able to correctly choose between a giant green "DOWNLOAD NOW!" advertisement/virus button and a smaller link to the actual desired file?

Post reply on HN