Live data from Hacker News

Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

anthropic.com

611–620 of 758 posts

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#611
post #605
post #589

Earlier quoted context omitted.

I will repeat my question from one of the previous threads: Can someone explain these Aider benchmarks to me? They pass same 113 tests through llm every time. Why they then extrapolate ability of llm to pass these 113 basic python challenges to the general ability to produce/edit code? Couldn't LLM provider just fine-tune their model for these tasks specifically - since they are static - to get ad value? Did anyone e…

> Couldn't LLM provider just fine-tune their model for these tasks specifically - since they are static - to get ad value? They could. They would easily be found out as they loose in real world usage or improved new unique benchmarks. If you were in charge of a large and well funded model, would you rather pay people to find and "cheat" on LLM benchmarks by training on them, or would you pay people to identify benchm…

> If you were in charge of a large and well funded model, would you rather pay people to find and "cheat" on LLM benchmarks by training on them, or would you pay people to identify benchmarks and make reasonably sure they specifically get excluded from training data?

You should already know by now that economic incentives are not always aligned with science/knowledge...

This is the true alignment problem, not the AI alignment one hahaha

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#612
post #589

The new Sonnet tops aider's code editing leaderboard at 84.2%. Using aider's "architect" mode it sets the SOTA at 85.7% (with DeepSeek as the "editor" model). 84% Claude 3.5 Sonnet 10/22 80% o1-preview 77% Claude 3.5 Sonnet 06/20 72% DeepSeek V2.5 72% GPT-4o 08/06 71% o1-mini 68% Claude 3 Opus It also sets SOTA on aider's more demanding refactoring benchmark with a score of 92.1%! 92% Sonnet 10/22 75% o1-preview 72%…

I will repeat my question from one of the previous threads: Can someone explain these Aider benchmarks to me? They pass same 113 tests through llm every time. Why they then extrapolate ability of llm to pass these 113 basic python challenges to the general ability to produce/edit code? Couldn't LLM provider just fine-tune their model for these tasks specifically - since they are static - to get ad value? Did anyone e…

There is an opportunity to develop black-box benchmarks and offer them to LLM providers to support their testing phase. If I were in their place, I would find it incredibly valuable to have such tamper-proof testing before releasing a model.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#613
post #605
post #589

Earlier quoted context omitted.

I will repeat my question from one of the previous threads: Can someone explain these Aider benchmarks to me? They pass same 113 tests through llm every time. Why they then extrapolate ability of llm to pass these 113 basic python challenges to the general ability to produce/edit code? Couldn't LLM provider just fine-tune their model for these tasks specifically - since they are static - to get ad value? Did anyone e…

> Couldn't LLM provider just fine-tune their model for these tasks specifically - since they are static - to get ad value? They could. They would easily be found out as they loose in real world usage or improved new unique benchmarks. If you were in charge of a large and well funded model, would you rather pay people to find and "cheat" on LLM benchmarks by training on them, or would you pay people to identify benchm…

They cannot be found out as long as there is no better evaluation. Sure, if they produce obvious nonsense, but the point of a systematic evaluation is exactly to overcome subjective impressions based on individual examples as a notion of quality.

Also, you are right that excluding test data from the training data improves your model. However, given the insane amounts of training data, this requires significant effort. If that additionally leads to your model performing worse in existing leaderboards, I doubt that (commercial) organizations would pay for such an effort.

And again, as long as there is no better evaluation method, you still won't know how much it really helps.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#614

Earlier quoted context omitted.

Because its a finetune of 3.5 optimized for the use case of computer use. Its actually accurate and its not a 3.6.

I don't think that's correct. This looks like a new model. Significant jump in math and gpqa scores.

I would assume that 3.5 means that the base training (which takes weeks/month) wasn't changed but only finetuning happened.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#615

Earlier quoted context omitted.

> There is nothing magical or special about human brain. There is a lot about the human brain that even the world's top neuroscientists don't know. There's plenty of magic about it if we define magic as undiscovered knowledge. There's also no consensus among top AI researchers that current techniques like LLMs will get us anywhere close to AGI. Nothing I've seen on current models (not even o1-preview) suggests to me…

Defining AGI as “can reason about 5MLOC” is ridiculous. When do the goal posts stop moving? When a computer can solve time travel? Babies have behavior all the time that is no more differentiable from what an LLM does on a normal basis (including terrible logic and hallucinations). The majority of people on the planet can barely reason about how any given politician will affect them, even when there’s a billion resou…

Babies can at least manipulate the physical world. Large language model can never be defined as AGI until it can control a general purpose robot, similar to how human brain controls our body's motor functions.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#616
post #344

This is actually a huge deal. As someone building AI SaaS products, I used to have the position that directly integrating with APIs is going to get us most of the way there in terms of complete AI automation. I wanted to take at stab at this problem and started researching some daily busineses and how they use software. My brother-in-law (who is a doctor) showed me the bespoke software they use in his practice. Runni…

I'm a bit skeptical about this working well enough to handle exceptions as soon as something out of the ordinary occurs. But it seems this could work great for automated testing.

Has anyone tried asking "use computer" to do "Please write a selenium/capybara/whatever test for filling out this form and sending it?"

That would take away some serious drudge work. And it's not a big problem if it fails, contrary to when it makes a mistake in filling out a form in an actual business process.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#617
post #605
post #589

Earlier quoted context omitted.

I will repeat my question from one of the previous threads: Can someone explain these Aider benchmarks to me? They pass same 113 tests through llm every time. Why they then extrapolate ability of llm to pass these 113 basic python challenges to the general ability to produce/edit code? Couldn't LLM provider just fine-tune their model for these tasks specifically - since they are static - to get ad value? Did anyone e…

> Couldn't LLM provider just fine-tune their model for these tasks specifically - since they are static - to get ad value? They could. They would easily be found out as they loose in real world usage or improved new unique benchmarks. If you were in charge of a large and well funded model, would you rather pay people to find and "cheat" on LLM benchmarks by training on them, or would you pay people to identify benchm…

This market is all about hype and mindshare, proper testing is hard and not performed by individuals, so there are no incentives not to train a bit on the test set.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#618

Earlier quoted context omitted.

Every time I see this argument made, there seems to be a level of complexity and/or operational cost above which people throw up their hands and say "well of course we can't do that". I feel like we will see that again here as well. It really is similar to the self-driving problem.

This is because it will be absolutely catastrophic economically when the majority of high paying jobs can be automated and owned by a few billionaires. Then what will go along with this catastrophe will be all the service people who had jobs to support the people with high paid jobs, they're fucked too. People don't want to have to face that. We'd be losing access to food, shelter, insurance, purpose. I can't blame p…

If most people are unemployed, modern capitalism as we know it will collapse. I'm not sure that's in the interests of the billionaires. Perhaps some kind of a social safety net will be implemented.

But I do agree, there is no reason to be enthusiastic about any progress in AI, when the goal is simply automating people's jobs away.

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#619
post #611
post #605

Earlier quoted context omitted.

> Couldn't LLM provider just fine-tune their model for these tasks specifically - since they are static - to get ad value? They could. They would easily be found out as they loose in real world usage or improved new unique benchmarks. If you were in charge of a large and well funded model, would you rather pay people to find and "cheat" on LLM benchmarks by training on them, or would you pay people to identify benchm…

> If you were in charge of a large and well funded model, would you rather pay people to find and "cheat" on LLM benchmarks by training on them, or would you pay people to identify benchmarks and make reasonably sure they specifically get excluded from training data? You should already know by now that economic incentives are not always aligned with science/knowledge... This is the true alignment problem, not the AI…

The AI alignement problem and the people alignment problem are actually the same problem! :D

One is just a bit harder due to the less familiar mind "design".

Re: Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku

#620

The new Sonnet tops aider's code editing leaderboard at 84.2%. Using aider's "architect" mode it sets the SOTA at 85.7% (with DeepSeek as the "editor" model). 84% Claude 3.5 Sonnet 10/22 80% o1-preview 77% Claude 3.5 Sonnet 06/20 72% DeepSeek V2.5 72% GPT-4o 08/06 71% o1-mini 68% Claude 3 Opus It also sets SOTA on aider's more demanding refactoring benchmark with a score of 92.1%! 92% Sonnet 10/22 75% o1-preview 72%…

[deleted]
Post reply on HN