Live data from Hacker News

Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

anthropic.com

361–370 of 758 posts

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#362

Earlier quoted context omitted.

This is, craaaaaazzzzzy. I'm just a layman, but to me, this is the most compelling evidence that things are starting to tilt toward AGI that I've ever seen.

Nah, it's the equivalent of seeing faces in static, or animals in clouds. Our brains are hardwired to see patterns, even when there are none. A similar, and related, behavior is seeing intent and intelligence in random phenomenon.

So it's behaving like our brains. Yet it's not AGI.

Does that mean our brains do not implement General Intelligence?

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#363
post #344

This is actually a huge deal. As someone building AI SaaS products, I used to have the position that directly integrating with APIs is going to get us most of the way there in terms of complete AI automation. I wanted to take at stab at this problem and started researching some daily busineses and how they use software. My brother-in-law (who is a doctor) showed me the bespoke software they use in his practice. Runni…

Absolutely! This reminds me of the humanoid robots vs specialized machines debate.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#364

This needs more discussion: Claude using Claude on a computer for coding https://youtu.be/vH2f7cjXjKI?si=Tw7rBPGsavzb-LNo (3 mins) True end-user programming and product manager programming are coming, probably pretty soon. Not the same thing, but Midjourney went from v.1 to v.6 in less than 2 years. If something similar happens, most jobs that could be done remotely will be automatable in a few years.

Every time I see this argument made, there seems to be a level of complexity and/or operational cost above which people throw up their hands and say "well of course we can't do that". I feel like we will see that again here as well. It really is similar to the self-driving problem.

I feel pain for the people who will be employed to "prompt engineer" the behavior of these things. When they inevitably hallucinate some insane behavior a human will have to take blame for why it's not working.. and yea, that'll be fun to be on the receiving end of.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#365

While I was initially impressed with it's context window, I got so sick of fighting with Claude about what it was allowed to answer I quit my subscription after 3 months. Their whole policing AI models stance is commendable but ultimately renders their tools useless. It actually started arguing with me about whether it was allowed to help implement a github repository's code as it might be copywritten... it was MIT l…

I just include text that I own the device in question and that I have a legal team watching my every move. It's stupid, I agree, but not insurmountable. I had less refusals with Claude 3 Opus.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#367

It improves to 25.9 over the previous version of Claude 3.5 Sonnet (24.4) on NYT Connections: https://github.com/lechmazur/nyt-connections/ .

Perhaps it's just because English is not my native language, but the prompt 3 isn't quite clear at the beginning when it says "group of four. Words (...)". It is not explained what the group of four must be, if I add to the prompt "group of four words" Claude 3.5 manages to answer it, while without it, Claude tells it is not that clear and can't answer

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#368
post #160
post #9

I still feel like the difference between Sonnet and Opus is a bit unclear. Somewhere on Anthropic's website it says that Opus is the most advanced, but on other parts it says Sonnet is the most advanced and also the fastest. The UI doesn't make the distinction clear either. Then on Perplexity, Perplexity says that Opus is the most advanced, compared to Sonnet. And finally, in the table in the blogpost, Opus isn't eve…

By reputation -- I can't vouch for this personally, and I don't know if it'll still be true with this update -- Opus is still often better for things like creative writing and conversations about emotional or political topics.

Yes, (old) 3.5 Sonnet is distinctly worse at emotional intelligence, flexibility, expressiveness and poetry.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#369
I really don't get their model. They have very advanced models, but the service overall seems to be a jumble of priorities. Some examples:

Anthropic doesn't offer an unlimited chatbot service, only plans that give you "more" usage, whatever that means. If you have an API key, you are "unlimited," so they have the capability. Why doesn't the chatbot allow one to use their API key in the Claude app to get unlimited usage? (Yes, I know there are third-party BYOK tools. That's not the question.)

Claude appears to be smart enough to make an Excel spreadsheet with simple formulae. However, it is apparently prevented from making any kind of file. Why? What principle underlies that guardrail that does not also apply to Computer Use?

Really want to make Claude my daily driver, but right now it often feels too much like a research project.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#370

And today I realized that despite it being an extremely common activity, we don’t really have a word for “using the computer” which is distinct from “computing”. It’s funny because AI models are always “using a computer” but now they can “use your computer.”

Computering

[deleted]
Post reply on HN