Live data from Hacker News

Claude 3.5 Sonnet

anthropic.com

281–287 of 287 posts

Re: Claude 3.5 Sonnet

#281

Which is the goto leaderboard for determining which AI model is best for for answering devops / computer science questions / generating code? Wondering where Claude falls on this. Recently canceled openai subscription because too much lag and crashes. Switched to Gemini because their webinterface is faster and rock solid. Makes me think the openai backend and frontend engineers don't know what they are doing compared…

One extensive benchmark I like is https://bigcode-bench.github.io/

It places Claude 3.5 Sonnet in third position.

Re: Claude 3.5 Sonnet

#283

Earlier quoted context omitted.

Funny anecdote for you. I usually test LLM's by attempting to play DnD 5e with them. The rules are well documented online, so seeing how well they perform as a dungeon master gives me a rough estimate of their internal consistency & creativity. For this, Claude performs fantastically. Outperforms every other LLM I've tested by a wide margin. However, when (as a player character) I tried to convince an NPC trickster m…

What prompts do you use for DnD / dungeon master? Think this would be great for solo campaigns.

Claude didn't require a whole lot of prompt wrangling to get started (also part of the test). Just talk to it like you would normally ("Hey, you know the DnD 5e rules? Could you make me a character sheet to fill out? Ready to play?" etc.)

Re: Claude 3.5 Sonnet

#284

Opus remained better than GPT for me, even after the release of GPT-4o. VERY happy to see an even further improvement beyond that, Claude is a terrific product and given the news that GPT-5 only began its training several weeks ago I don't see any situation where Anthropic is dethroned in the near term. There are only two parts of Anthropic's offering I'm not a fan of: - Lack of conversation sharing: I had a conversa…

Both GPT-4 and 4o have been completely useless for coding in the past couple of weeks for me - constant errors, and not just your typical LLM inaccuracies but incapable of producing a few lines of self-consistent code e.g. defines variables foo on one line and refers to it as bar on the next, or it misspells it as foox.

The level of misspelling is insane at the moment. It does it almost 50%+ of the times. I just started using claude 3.5 and the difference is night and day.

Re: Claude 3.5 Sonnet

#286

Awesome, can’t wait to try this. I wish the big AI labs would make more frequent model improvements, like on a monthly cadence, as they continue to train and improve stuff. Also seems like a good way to do A/B testing to see which models people prefer in practice.

That's a hard ask given the actual training runs are multi-month, and have distinguished pretraining and refinement phases.

True, but they could have multiple parallel running training processes going on at the same time. And they could release models that result from partial training checkpoints if they can quantify that they are better than the last released model and also "safe."

Re: Claude 3.5 Sonnet

#287

Earlier quoted context omitted.

https://openai.com/index/openai-board-forms-safety-and-secur... (May 28th) > OpenAI has recently begun training its next frontier model and we anticipate the resulting systems to bring us to the next level of capabilities on our path to AGI.

No doubt openai have been training big models for the last year. If “gpt5” is only just starting it means recent training runs have had disappointing results and have been passed off as “Gpt4o” or whatever. The value of all the AI companies is predicated on high chance of AGI, and gpt5 failing to be revolutionary may pop the whole bubble (+10 trillion of market cap)

> If “gpt5” is only just starting it means recent training runs have had disappointing results and have been passed off as “Gpt4o” or whatever.

Sora probably took a lot of cluster time don't you think?

Post reply on HN