Live data from Hacker News

Claude 3.5 Sonnet

anthropic.com

221–230 of 287 posts

Re: Claude 3.5 Sonnet

#222

Opus remained better than GPT for me, even after the release of GPT-4o. VERY happy to see an even further improvement beyond that, Claude is a terrific product and given the news that GPT-5 only began its training several weeks ago I don't see any situation where Anthropic is dethroned in the near term. There are only two parts of Anthropic's offering I'm not a fan of: - Lack of conversation sharing: I had a conversa…

I do wonder if GPT quality fluctuates seasonally, or with electricity costs, in an engineering effort to balance costs with performance. I agree on all your points, but would like to emphasize that I really do enjoy the voice input voice output thing that chatgpt's app has. Its not how I use it when working, but when commuting, a lot of times, I'll turn on the the chatgpt app and have a conversation with it exploring…

Short of switching between models (which at least OpenAI definitely does for free customers, but I believe they always indicate it), how would that work? Different quantizations?

Re: Claude 3.5 Sonnet

#223

This is amazing - I far prefer the personality of Claude to GPT-4 series models. Also, with coding tasks, Claude-3-Opus and been far better for me vs gpt-4-turbo and gpt-4o both. Looking forward to giving it a spin. Seems like it's doing better than GPT-4o in most benchmarks though I'd like to see if its speed is comparable or not. Also, eagerly awaiting the LMSYS blind comparison results!

>I far prefer the personality of Claude to GPT-4 series models.

This new Sonnet seems way less human-like than even old Sonnet, let alone Opus. It's practically devoid of character. It's smart, though.

Re: Claude 3.5 Sonnet

#224

On a first glance, CS3.5 appears to be slightly faster than gpt-4o (62 vs 49 tok/sec) and slightlhy less capable (78% vs 89% accuracy on our internal reasoning benchmark). When initially launched, gpt-4o had speed of over 100 tok/sec, surprised that speed went down as fast.

I'm not asking for actual examples, but what kind of thing is in your internal reasoning benchmark?

Things like “summarize this text in exactly 14 words”, programming questions, unstructured data to structured data transformations and so on…

Re: Claude 3.5 Sonnet

#225

Earlier quoted context omitted.

I'm not asking for actual examples, but what kind of thing is in your internal reasoning benchmark?

Things like “summarize this text in exactly 14 words”, programming questions, unstructured data to structured data transformations and so on…

Do you let it use CoT? I think that first one is pretty hard if you have to produce it directly one token at a time, but I guess that's kind of the point.

Re: Claude 3.5 Sonnet

#226

Opus remained better than GPT for me, even after the release of GPT-4o. VERY happy to see an even further improvement beyond that, Claude is a terrific product and given the news that GPT-5 only began its training several weeks ago I don't see any situation where Anthropic is dethroned in the near term. There are only two parts of Anthropic's offering I'm not a fan of: - Lack of conversation sharing: I had a conversa…

Both GPT-4 and 4o have been completely useless for coding in the past couple of weeks for me - constant errors, and not just your typical LLM inaccuracies but incapable of producing a few lines of self-consistent code e.g. defines variables foo on one line and refers to it as bar on the next, or it misspells it as foox.

I've been experiencing bizarre typos and misspellings that I've come to describe as the model being drunk. Things like it writing peremeter instead of parameter

Re: Claude 3.5 Sonnet

#227

Earlier quoted context omitted.

No doubt openai have been training big models for the last year. If “gpt5” is only just starting it means recent training runs have had disappointing results and have been passed off as “Gpt4o” or whatever. The value of all the AI companies is predicated on high chance of AGI, and gpt5 failing to be revolutionary may pop the whole bubble (+10 trillion of market cap)

Sam said on Lex's podcast that people should temper their expectations for GPT-5, not in that it will necessarily suck, but that they want to ramp up ability slowly over time rather than discrete large steps.

Sounds like an excuse tbh. Esp when other companies are pushing ahead beyond OAI and open source is close to rivaling them

Re: Claude 3.5 Sonnet

#228
post #163

Earlier quoted context omitted.

I also believe that gpt-4o was originally called gpt-5. If you look at the image generation on their website from gpt-4o which has not been released, I believe that along with the voice caused Ilya to declare mission accomplished (AGI) and that is why there was a coup. The coup failed because no one wanted to wrap up the company or change the way it operated because they would lose a lot of money. The reason the name…

I don't hate this speculation, I just don't buy it at all. 4o's about the same in terms of reasoning as 4. People don't find the text abilities that much more usable over 4 (at least on the LMS leaderboard). It's faster and has audio2audio capabilities alongside new native image stuff I think, but how exactly is that AGI if 4 isn't? These models understanding and reasoning ability is still far too weak to do any seri…

Scroll to Explorations of Capabilities: https://openai.com/index/hello-gpt-4o/

That combined with the voice was probably considered AGI by Ilya.

Re: Claude 3.5 Sonnet

#229
post #202

Earlier quoted context omitted.

Personal prompting style, I imagine,

People really, really , underestimate how important prompting is. I would be confident in stating that half the people who complain about a model are actually just suffering from poor prompting.

And what makes you so confident that all those people are using different prompt styles when comparing models? You think most people don’t even understand the bare basics of how to compare two products?

Re: Claude 3.5 Sonnet

#230

Earlier quoted context omitted.

People really, really , underestimate how important prompting is. I would be confident in stating that half the people who complain about a model are actually just suffering from poor prompting.

And what makes you so confident that all those people are using different prompt styles when comparing models? You think most people don’t even understand the bare basics of how to compare two products?

That's the point: maybe someone has a personal prompting style that works great with Claude but gives worse results with GPT-4.

They might complain that GPT-4 is rubbish in comparison to Claude, but someone with a different personal prompting style might experience the opposite.

Post reply on HN