Live data from Hacker News

Claude 3.5 Sonnet

anthropic.com

251–260 of 287 posts

Re: Claude 3.5 Sonnet

#251

This is amazing - I far prefer the personality of Claude to GPT-4 series models. Also, with coding tasks, Claude-3-Opus and been far better for me vs gpt-4-turbo and gpt-4o both. Looking forward to giving it a spin. Seems like it's doing better than GPT-4o in most benchmarks though I'd like to see if its speed is comparable or not. Also, eagerly awaiting the LMSYS blind comparison results!

I'm surprised there isn't a single mention of Gemini 1.5 Pro. I've been using it for about a month because it came for free with my Google setup and I've been pretty happy. Not for coding but mostly for business tasks like writing minutes from transcripts, summarizing long legal documents,... and the long context length has been awesome. It also conveniently integrates with the rest of my google setup like Drive.

IIRC it also ranked only behind gpt4o on benchmarks.

Re: Claude 3.5 Sonnet

#252
post #230

Earlier quoted context omitted.

And what makes you so confident that all those people are using different prompt styles when comparing models? You think most people don’t even understand the bare basics of how to compare two products?

That's the point: maybe someone has a personal prompting style that works great with Claude but gives worse results with GPT-4. They might complain that GPT-4 is rubbish in comparison to Claude, but someone with a different personal prompting style might experience the opposite.

Having a prompting style that works with a model but not quite with another is much different than "suffering from poor prompting" the previous person was accusing others.

And given that those are tools, it's more like "the model can work with the user's prompts" rather than "the user's prompts are adapted to the model".

Unless we're here for an ego trip.

Re: Claude 3.5 Sonnet

#253

For Anthropic devs out there: Please consider adopting a naming convention that will automatically upgrade API users to the latest version when available. Eg. there should be just 'claude-sonnet'.

But that would be similar to using `latest` tag with docker images, which is rarely a good idea…

Re: Claude 3.5 Sonnet

#254

I have been so impressed with Claude. I routinely turn to it over ChatGPT. Has anyone else felt like 4o was a downgrade from 4? Anyway…this is exciting

Yeah, I had an impression that 4o, while much faster, is a downgrade. For me it often starts repeating things over and over when I start questioning things.

Re: Claude 3.5 Sonnet

#256
post #251

This is amazing - I far prefer the personality of Claude to GPT-4 series models. Also, with coding tasks, Claude-3-Opus and been far better for me vs gpt-4-turbo and gpt-4o both. Looking forward to giving it a spin. Seems like it's doing better than GPT-4o in most benchmarks though I'd like to see if its speed is comparable or not. Also, eagerly awaiting the LMSYS blind comparison results!

I'm surprised there isn't a single mention of Gemini 1.5 Pro. I've been using it for about a month because it came for free with my Google setup and I've been pretty happy. Not for coding but mostly for business tasks like writing minutes from transcripts, summarizing long legal documents,... and the long context length has been awesome. It also conveniently integrates with the rest of my google setup like Drive. IIR…

Gemini in general is terrible. Way too many mistakes. If you use it via the API it repeats itself constantly. At least it's the model that is the easiest to jailbreak and will happiliy give you a tutorial on how to make a bomb if you ask politely ;) Very ironic considering how Google emphasizes "safety".

Re: Claude 3.5 Sonnet

#257

Opus remained better than GPT for me, even after the release of GPT-4o. VERY happy to see an even further improvement beyond that, Claude is a terrific product and given the news that GPT-5 only began its training several weeks ago I don't see any situation where Anthropic is dethroned in the near term. There are only two parts of Anthropic's offering I'm not a fan of: - Lack of conversation sharing: I had a conversa…

> I had a conversation with Claude where I asked it to reverse engineer some assembly code and it did it perfectly on the first try. I was stunned

I share the same experience with you but with Claude 3 Sonnet. I can’t count how many times I’ve shared some code with Claude with barely any hope because other GPTs failed aswell, yet, Claude surprised me and performed the task with success.

I’ve actually reached to the point that I expressed my gratitude to Claude because of how well it performs on coding tasks and other tasks in general. I don’t know what Anthropic did, but something did they right.

Being able to handle large amounts of tokens, “understand” and perform tasks on it & spit out large amounts of data back with barely any cut-offs (unlike Gemini) has made me feel like Claude is at the moment the best option.

Re: Claude 3.5 Sonnet

#258

Doesn't look to be available on Bedrock yet. Maybe tomorrow, since the article says June 21st also? We truly live in the future...

It's available now! Sorry for the delay.

I don't see it yet on us-west-2 (bedrock -> model access). Do you mean another region? Or does the account have to be in a special canary group?

Re: Claude 3.5 Sonnet

#259
post #163

Earlier quoted context omitted.

What I understand is that it's GPT 6 that just went into training, and that GPT 5 is complete and being delayed until after the U.S. election.

I also believe that gpt-4o was originally called gpt-5. If you look at the image generation on their website from gpt-4o which has not been released, I believe that along with the voice caused Ilya to declare mission accomplished (AGI) and that is why there was a coup. The coup failed because no one wanted to wrap up the company or change the way it operated because they would lose a lot of money. The reason the name…

It's speculation with no basis at all, OAI has a track record of releasing half step models and 4o is no different just like 3 to 3.5 and the numerous subsequent 3.5 releases.

If you've used 4 and 4o they are too similar for 4o to have been trained from scratch

Re: Claude 3.5 Sonnet

#260
post #45

Might look small, but the needle in a haystack numbers they report in the model card addenda at 200k are also a massive improvement towards “Proving a negative”… I.e. your answer does not exist in your text. %99.7 vs 98.3 for Opus https://cdn.sanity.io/files/4zrzovbb/website/fed9cc193a14b84...

Could you explain how these two are related? That benchmark seems to be asking for very specific information inside a large body of text. For LLMs, that seems quite a different task compared to proving a negative. Any improvements on proving a negative would mean less hallucinations and would be a huge deal.
Post reply on HN