Computer use seems it might be good for e2e tests.
Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
31–40 of 758 posts
Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
#32I have been a paying ChatGPT customer for a long time (since the very beginning). Last week I've compared ChatGPT to Claude and the results (to my eye) were better, the output better structured and the canvas works better. I'm on the edge of jumping ship.
o1 is pretty decent as a rotor rooter, ie the type of task that requires both lots of instruction as well as lots of context. I honestly think it works half as well as it does now because it’s able to properly mull through the true intent of the user that usually takes the multiple shots that nobody has the patience to do.
Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
#33Why not rev the numbers? "3.5" vs. "3.5 New" feels weird -- is there a particular reason why Anthropic doesn't want to call this 3.6 (or even 3.5.1)?
Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
#34Looks like it just takes a screenshot and can't scroll so it might miss things. Claude 3.5 Haiku will be released later this month.
It can actually scroll.
Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
#35I still feel like the difference between Sonnet and Opus is a bit unclear. Somewhere on Anthropic's website it says that Opus is the most advanced, but on other parts it says Sonnet is the most advanced and also the fastest. The UI doesn't make the distinction clear either. Then on Perplexity, Perplexity says that Opus is the most advanced, compared to Sonnet. And finally, in the table in the blogpost, Opus isn't eve…
I think they originally announced that Opus would get a 3.5 update, but with every product update they are doing I'm doubting it more and more. It seems like their strategy is to beat the competition on a smaller model that they can train/tune more nimbly and pair it with outside-the-model product features, and it honestly seems to be working.
Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
#36Nice improvements in scores across the board, e.g.
> On coding, it [the new Sonnet 3.5] improves performance on SWE-bench Verified from 33.4% to 49.0%, scoring higher than all publicly available models—including reasoning models like OpenAI o1-preview and specialized systems designed for agentic coding.
I've been using Sonnet 3.5 for most of my AI-assisted coding and I'm already very happy (using it with the Zed editor, I love the "raw" UX of its AI assistant), so any improvements, especially seemingly large ones like this are very welcome!
I'm still extremely curious about how Sonnet 3.5 itself, and its new iteration are built and differ from the original Sonnet. I wonder if it's in any way based on their previous work[0] which they used to make golden-gate Claude.
[0]: https://transformer-circuits.pub/2024/scaling-monosemanticit...
Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
#37From the computer use video demo, that's a lot of API calls. Even though Claude 3.5 Sonnet is relatively cheap for its performance, I suspect computer use won't be. It's a very good idea that Anthropic upfront that it isn't perfect. And it's guaranteed that there will be a viral story where Claude will accidentally delete something important with it. I'm more interested in Claude 3.5 Haiku, particularly if it is inde…
Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
#38A comment about the video: Sam Runger talks wayyy too fast, in particular at the beginning.
Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
#39Pretty cool for sure.
Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
#40Earlier quoted context omitted.
Opus has been stuck on 3.0, so Sonnet 3.5 is better for most things as well as cheaper.
> Opus has been stuck on 3.0, so Sonnet 3.5 is better So for example, Perplexity is wrong here implying that Opus is better than Sonnet? https://i.imgur.com/N58I4PC.png