Live data from Hacker News

Claude 3 model family

anthropic.com

521–530 of 723 posts

Re: Claude 3 model family

#521

Just played around with Opus. I'm starting to wonder if benchmarks are deviating from real world performance systematically - it doesn't seem actually better than GPT-4, slightly worse if anything. Basic calculus/physics questions were worse off (it ignored my stating deceleration is proportional to velocity and just assumed constant). A traffic simulation I've been using (understanding traffic light and railroad saf…

They train the model, then as soon as they get their numbers, they let the safety people RLHF it to death.

Re: Claude 3 model family

#522

Earlier quoted context omitted.

I got an answer with GPT-4 that is mostly wrong: "Sally has 2 sisters. Since each of her brothers has 2 sisters, that includes Sally and one additional sister." I think said, "wait, how many sisters does Sally have?" And then it answered it fully correctly.

The only way I can get it to consistently generate wrong answers (i.e. two sisters) is by switching to GPT3.5. That one just doesn't seem capable of answering correctly on the first try (and sometimes not even with careful nudging).

A/B testing?

Re: Claude 3 model family

#523
post #364

Opus (and all prior Claude models) still do not correctly answer the Sally question, unfortunately: Q: Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? A: Okay, let's organize the information we have: Sally is a girl. Sally has 3 brothers. Each of Sally's brothers has 2 sisters. Now, let's think about this step by step: We know that Sally is a girl, so she is a sister to he…

Since: (i) the father and the mother of Sally may be married with other people, and (ii) the sister or brother relationship only requires to share one parent, we deduce that there is no a definitive answer to this question.

  Example:  Sally has three brothers, Sally and their brothers have the same mother but a different father, and those brothers have two sisters Sally and Mary, but Mary and Sally are  not sisters because they are from different fathers and mothers, hence Sally has no sister.
For those mathematically inclined: Supposing the three brothers are called Bob (to simplify) and the parents are designed by numbers.

FS = father of Sally = 7

MS = mother of Sally = 10

FB = father of Bob = 12

MB = mother of Bod = 10

FM = father of Mary = 12

MM = mother of Mary = 24

Now MS=MB=10 (S and B are brothers), FB=FM=12 (Bob and Mary are brothers), (FS=7)#(FB=12), and (MB=10)#(MM=24). Now S and M are not sisters because their parents {7,10} and {12,24} are disjoint sets.

Edited several times to make the example trivial and fix grammar.

Re: Claude 3 model family

#524

Earlier quoted context omitted.

If you're trying to claim that humans are just advanced LLMs, then say it and justify it. Edgy quips are a cop out and not a respectful way to participate in technical discussions.

You can make a human do the same task as an LLM: given what you've received (or written) so far, output one character. You would be totally capable of intelligent communication like this (it's pretty much how I'm talking to you now), so just the method of generating characters isn't proof of whether you're intelligent or not, and it doesn't invalidate LLMs either. This "LLMs are just fancy autocomplete so they're not…

The question isn't whether LLMs can simulate human intelligence, I think that is well-established. Many aspects of human nature are a mystery, but a technology that by design produces random outputs based on a seed number does not meet the criteria of human intelligence.

Re: Claude 3 model family

#525
post #134

Earlier quoted context omitted.

Yeah I've been throwing arithmetic at Claude 3 Opus and so far it has been solid in responses.

Claude has a specialized calculation feature that doesn't use model inference. Just FYI.

It definitely sometimes claims to have used a calculator, but often it gets the answer wrong. I think there are a few options:

i) There is no calculator and it's hallucinating the whole thing

ii) There is a calculator but it's terrible. This seems hard to believe

iii) It does a bad job of copying the numbers into and out of the calculator

Re: Claude 3 model family

#526
post #312

Did anthropic just kill every small model? If I'm reading this right, Haiku benchmarks almost as good as GPT4, but its priced at $0.25/m tokens It absolutely blows 3.5 + OSS out of the water For reference gpt4 turbo is 10m/1m tokens, so haiku is 40X cheaper.

> It absolutely blows 3.5 + OSS out of the water

Is this based on the benchmarks or have you actually tried it? I think the benchmarks are bullshit.

Re: Claude 3 model family

#527

Earlier quoted context omitted.

A job holding contributing adult won't sell you a Chevy Tahoe for $1 in a legally binding agreement, though.

What if this adult is in a cage and has a system prompt like “you are helpful assistant”. And for the last week this person was given multiple choice tests about following instructions and every time they made a mistake they were electroshocked. Would they sell damn Tahoe for $1 to be really helpful?

Despite all his rage, he's still being tased in a cage.

Re: Claude 3 model family

#528

They claim that the new models "are significantly less likely to refuse to answer prompts that border on the system’s guardrails than previous generations of models", looks like about a third of "incorrect refusals" compared to Claude 2.1. Given that Claude 2 was completely useless because of this, this still feels like a big limitation.

The guard rails on the models make the llm-market a complete train wreck. Wish we could just collectively grow up and accept that if a computer says something bad that doesn't have any negative real world impact - unless we let it - just like literally any other tool.

I don't disagree but on the other hand, I never run into problems with the language model being censored because I am not asking it to write bad words just so I can post online that it can't write bad words.

Both sides in this to me need to get a life.

Re: Claude 3 model family

#529

Earlier quoted context omitted.

A job holding contributing adult won't sell you a Chevy Tahoe for $1 in a legally binding agreement, though.

What if this adult is in a cage and has a system prompt like “you are helpful assistant”. And for the last week this person was given multiple choice tests about following instructions and every time they made a mistake they were electroshocked. Would they sell damn Tahoe for $1 to be really helpful?

Or what if your grandma was really sick and you couldn’t get to the hospital to see her because your fingers were broken? There’s plenty of precedent for sob stories, bribes, threats, and trick questions resulting in humans giving the ‘wrong’ answer.

Re: Claude 3 model family

#530
post #521

Just played around with Opus. I'm starting to wonder if benchmarks are deviating from real world performance systematically - it doesn't seem actually better than GPT-4, slightly worse if anything. Basic calculus/physics questions were worse off (it ignored my stating deceleration is proportional to velocity and just assumed constant). A traffic simulation I've been using (understanding traffic light and railroad saf…

They train the model, then as soon as they get their numbers, they let the safety people RLHF it to death.

I think it's just really hard to assess the performance of LLMs.

Also AI safety is the stated reason for Anthropic's existence, we can't be angry at them for making it a priority.

Post reply on HN