Live data from Hacker News

Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

anthropic.com

751–758 of 758 posts

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#751

I've been saying this is coming for a long time, but my really smart SWE friend who is nevertheless not in the AI/ML space dismissed it as a stupid roundabout way of doing things. That software should just talk via APIs. No matter how much I argued regarding legacy software/websites and how much functionality is really only available through GUI, it seems some people are really put off by this type of approach. To me…

I recall 90's Macs had a 3rd party app that offered to observe your mouse/keyboard then automatically recommend routine tasks for you. As a young person I found that fascinating. It's interesting to see history renew itself.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#752

Earlier quoted context omitted.

Without people providing their prompts, it's impossible to say whether they are skilled or not, and their complaints or claims of "it worked with this prompt" without the output are also not possible to validate. Maybe there's a clue in there as to why these experiences seem so different. I'm glad GPTs don't get frustrated.

Ive spent thousands of hours, literally, learning the ropes, and continue to hone it. There is a much higher skill ceiling for prompting than there was for Google-fu.

Literally ropes as in RoPE, rotary positional embeddings?

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#753
post #344

This is actually a huge deal. As someone building AI SaaS products, I used to have the position that directly integrating with APIs is going to get us most of the way there in terms of complete AI automation. I wanted to take at stab at this problem and started researching some daily busineses and how they use software. My brother-in-law (who is a doctor) showed me the bespoke software they use in his practice. Runni…

Talking about ancient Windows software... Windows used to have an API for automation in the 2000s (I don't know if it still does). I wrote this MS Access script that ran and moved the cursor at exactly the pixel coordinates where buttons and fields were positioned in a GUI that we wanted to extract data from, in one of my first jobs. My boss used to do this manually. After a week he had millions of records ready to q…

PowerShell has some amazing capabilities.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#754
post #589

The new Sonnet tops aider's code editing leaderboard at 84.2%. Using aider's "architect" mode it sets the SOTA at 85.7% (with DeepSeek as the "editor" model). 84% Claude 3.5 Sonnet 10/22 80% o1-preview 77% Claude 3.5 Sonnet 06/20 72% DeepSeek V2.5 72% GPT-4o 08/06 71% o1-mini 68% Claude 3 Opus It also sets SOTA on aider's more demanding refactoring benchmark with a score of 92.1%! 92% Sonnet 10/22 75% o1-preview 72%…

I will repeat my question from one of the previous threads: Can someone explain these Aider benchmarks to me? They pass same 113 tests through llm every time. Why they then extrapolate ability of llm to pass these 113 basic python challenges to the general ability to produce/edit code? Couldn't LLM provider just fine-tune their model for these tasks specifically - since they are static - to get ad value? Did anyone e…

Conveniently, author of these benchmarks remains silent on topic every time. Think about it :)

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#755

Earlier quoted context omitted.

As generally intelligent beings, we can adapt to reading and producing 5M LOC, or to live in arctic climates, or to build a building in colonial or classical style as dictated by cost, taste, and other factors. That is generality in intelligence. I haven't moved any goal posts - it is your definition which is way too narrow.

You’re literally moving the goalposts right now. These models _are_ adapting to what you’re talking about. When Claude makes a model for haikus, how is that different than a poet who knows literally nothing about math but is fantastic at poetry? I’m sure as soon as Claude can handle 5MLOC you’ll say it should be 10, and it needs to make sure it can serve you a Michelin star dinner as well. That’s not AGI. Stop moving…

My point was it's not AGI, I don't even know what you're talking about or who you're replying to anymore.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#756
post #80

Seems like both: - AI Labs will eat some of the wrappers on top of their APIs - even complex ones like this. There are whole startups that are trying to build computer use. - AI is fitting _some_ scaling law - the best models are getting better and the "previously-state-of-the-art" models are fractions of what they cost a couple years ago. Though it remains to be seen if it's like Moore's Law or if incremental improv…

It seems a little silly to pretend there’s a scaling “law” without plotting any points or doing a projection. Without the mathiness, we could instead say that new models keep getting better and we don’t know how long that trend will continue.

"Law" might not be the right word - but there's no denying it's scaling with compute/data/model size. I suppose law happens after continued evidence over years.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#757

Earlier quoted context omitted.

Can you not use cross-region inference?

90% of our customers do not allow this due to data sovereignty. Bedrock here is lagging so far behind several customers assume AWS simply aren't investing here anymore - or if they are it's an afterthought - and a very expensive one at that. I've spoken with several account managers and SAs and they seem similarly frustrated with the continual response from above that useful models are "coming soon". You can't even B…

Hmm interesting, didn't realize that data sovereignty requirements were so stringent. Wonder how other cloud providers are doing in this sense considering GPU shortages across the board.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#758
post #221
post #92

Earlier quoted context omitted.

3ds doesn't have a version number to bump. Claude 3.5 does.

The 3 was the version number ;) Ds and ds lite were version 1 Dsi was 2 (as there was dsi software that didn’t run on ds or ds lite) And the 3ds was version 3.

https://www.perplexity.ai/search/what-is-the-meaning-of-the-...
Post reply on HN