Live data from Hacker News

Claude 3 model family

anthropic.com

421–430 of 723 posts

Re: Claude 3 model family

#421
post #364

Opus (and all prior Claude models) still do not correctly answer the Sally question, unfortunately: Q: Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? A: Okay, let's organize the information we have: Sally is a girl. Sally has 3 brothers. Each of Sally's brothers has 2 sisters. Now, let's think about this step by step: We know that Sally is a girl, so she is a sister to he…

I don't think this means much besides "It can't answer the Sally question".

Re: Claude 3 model family

#422

Could anyone recommend an open-source tool capable of simultaneously sending the same prompt to various language models like GPT-4, Gemini, and Claude, and displaying their responses side by side for comparison? I tried chathub in the past, but they decided to not release any more source as of now.

https://chat.lmsys.org/ Choose Arena (side-by-side), it has Claude 3 Opus, Sonnet and GPT-4

[deleted]

Re: Claude 3 model family

#423

Earlier quoted context omitted.

The reason we chucked loads of data at it was because we had no other options. If you wanted to write a function that classified a picture as a cat or a dog, good luck. With ML, you can learn such a function. That logic doesn’t extend to things we already know how to program computers to do. Arithmetic already works. We don’t need a neural net to also run the calculations or play a game of chess. We have specialized…

> We don’t need a neural net to also run the calculations or play a game of chess. That's actually one of the specific examples from the link I mentioned:- > In computer chess, the methods that defeated the world champion, Kasparov, in 1997, were based on massive, deep search. At the time, this was looked upon with dismay by the majority of computer-chess researchers who had pursued methods that leveraged human under…

What was considered “loads of compute” in 1998 is the kind of thing that can run on anyone’s phone today. Stockfish is extremely cheap compared with an LLM. Even a human-like model like Maia is tiny compared with even the smallest LLMs used these services.

Point is, LLM maximalists are wrong. Specialized software is better in many places. LLMs can fill in the gaps, but should hand off when necessary.

Re: Claude 3 model family

#424
post #364

Opus (and all prior Claude models) still do not correctly answer the Sally question, unfortunately: Q: Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? A: Okay, let's organize the information we have: Sally is a girl. Sally has 3 brothers. Each of Sally's brothers has 2 sisters. Now, let's think about this step by step: We know that Sally is a girl, so she is a sister to he…

[deleted]

Re: Claude 3 model family

#425

Could anyone recommend an open-source tool capable of simultaneously sending the same prompt to various language models like GPT-4, Gemini, and Claude, and displaying their responses side by side for comparison? I tried chathub in the past, but they decided to not release any more source as of now.

If you're willing to use the CLI, Simon Willison's llm library[0] should do the trick. [0] https://github.com/simonw/llm

I already have a cli client, but how to talk to multiple different LLMs at the same time? I guess I can script something with tmux.

Re: Claude 3 model family

#426
post #295

The APPS benchmark result of Claude 3 Opus at 70.2% indicates it might be quite useful for coding. The dataset measures the ability to convert problem descriptions to Python code. The average length of a problem is nearly 300 words. Interestingly, no other top models have published results on this benchmark. Claude 3 Model Card: https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bb... Table 1: Evaluation resul…

APPS has 3 subsets by difficulty level: introductory, interview, and competition. It isn't clear which subset Claude 3 was benchmarked on. Even if it is just "introductory" it is still pretty good, but it would be good to know.

Since they don’t state it, does it mean they tested it on the whole test set? If that’s the case, and we assume for simplicity that Opus solves all Intro problems and none of the Competition problems, it’d have solved 83%+ of the Interview level problems.

(There are 1000/3000/1000 problems in the test set in each level).

It’d be great if someone from Anthropic provides an answer though.

Re: Claude 3 model family

#427

Another naming disaster! Opus is better than sonnet? And sonnet is better than haiku? Perhaps this makes sense to people familiar with sonnets and haikus and opus....es? Nonsensical to me! I know everyone loves to hate on Google, but at least pro and ultra have a sort of sense of level of sophistication.

I think the intention was more "bigger" than better - but opus is an odd choice. haiku>sonnet>ballad maybe? haiku>sonnet>epic?

I don't know what an opus is, but the word sounds big. Maybe just because of the association with "Magnum Opus".

Haikus sound small, and sonnets kinda small too.

Re: Claude 3 model family

#428

Earlier quoted context omitted.

An AI smart enough to eclipse the average person on most basic tasks would even warrant far more hype than there is now.

Sure, but it would also be an IA much smarter than the ones we have now, because you cannot replace a human being with the current technology. You can augment one, making her perform the job of two or more humans before for some tasks, but you cannot replace them all, because the current tech cannot reasonably be used without supervision.

a lot of jobs are being replaced by AI already... comms/copywriting/customer service/off shored contract technicals roles especially.

Re: Claude 3 model family

#429
post #364

Opus (and all prior Claude models) still do not correctly answer the Sally question, unfortunately: Q: Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? A: Okay, let's organize the information we have: Sally is a girl. Sally has 3 brothers. Each of Sally's brothers has 2 sisters. Now, let's think about this step by step: We know that Sally is a girl, so she is a sister to he…

This is what my local LLama 13B Q5 had to say on this:

> There are two ways to approach this problem. First, we can count each person by gender and then add them up. Second, we can use the given information about the siblings and their genders to find a quicker solution.

> Let's go with the second method:

> 1. Sally has one sister (herself).

> 2. Each brother has two sisters. So, there are 3 sisters for each of her brothers.

> 3. Since there are three brothers, there are a total of 3 x 3 = 9 sisters for them all.

> 4. Now that we know Sally has one sister and the rest belong to her brothers, we can subtract her share from the total: 9 - 1 = 8 sisters.

>

> So, Sally has 8 sisters.

:|

Re: Claude 3 model family

#430

This is my highly advanced test image for vision understanding. Only GPT-4 gets it right some of the time - even Gemini Ultra fails consistently. Can someone who has access try it out with Opus? Just upload the image and say "explain the joke." https://i.imgur.com/H3oc2ZC.png

This is what I got on the Anthropic console, using Opus with temp=0:

> The image shows a cute brown and white bunny rabbit sitting next to a small white shoe or slipper. The text below the image says "He lost one of his white shoes during playtime, if you see it please let me know" followed by a laughing emoji.

> The joke is that the shoe does not actually belong to the bunny, as rabbits do not wear shoes. The caption is written as if the bunny lost its own shoe while playing, anthropomorphizing the rabbit in a humorous way. The silly idea of a bunny wearing and losing a shoe during playtime is what makes this a lighthearted, funny image.

Post reply on HN