Just played around with Opus. I'm starting to wonder if benchmarks are deviating from real world performance systematically - it doesn't seem actually better than GPT-4, slightly worse if anything. Basic calculus/physics questions were worse off (it ignored my stating deceleration is proportional to velocity and just assumed constant). A traffic simulation I've been using (understanding traffic light and railroad saf…
Claude 3 model family
521–530 of 723 posts
Re: Claude 3 model family
#522Earlier quoted context omitted.
I got an answer with GPT-4 that is mostly wrong: "Sally has 2 sisters. Since each of her brothers has 2 sisters, that includes Sally and one additional sister." I think said, "wait, how many sisters does Sally have?" And then it answered it fully correctly.
The only way I can get it to consistently generate wrong answers (i.e. two sisters) is by switching to GPT3.5. That one just doesn't seem capable of answering correctly on the first try (and sometimes not even with careful nudging).
Re: Claude 3 model family
#523Opus (and all prior Claude models) still do not correctly answer the Sally question, unfortunately: Q: Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? A: Okay, let's organize the information we have: Sally is a girl. Sally has 3 brothers. Each of Sally's brothers has 2 sisters. Now, let's think about this step by step: We know that Sally is a girl, so she is a sister to he…
Example: Sally has three brothers, Sally and their brothers have the same mother but a different father, and those brothers have two sisters Sally and Mary, but Mary and Sally are not sisters because they are from different fathers and mothers, hence Sally has no sister.
For those mathematically inclined: Supposing the three brothers are called Bob (to simplify) and the parents are designed by numbers.FS = father of Sally = 7
MS = mother of Sally = 10
FB = father of Bob = 12
MB = mother of Bod = 10
FM = father of Mary = 12
MM = mother of Mary = 24
Now MS=MB=10 (S and B are brothers), FB=FM=12 (Bob and Mary are brothers), (FS=7)#(FB=12), and (MB=10)#(MM=24). Now S and M are not sisters because their parents {7,10} and {12,24} are disjoint sets.
Edited several times to make the example trivial and fix grammar.
Re: Claude 3 model family
#524Earlier quoted context omitted.
If you're trying to claim that humans are just advanced LLMs, then say it and justify it. Edgy quips are a cop out and not a respectful way to participate in technical discussions.
You can make a human do the same task as an LLM: given what you've received (or written) so far, output one character. You would be totally capable of intelligent communication like this (it's pretty much how I'm talking to you now), so just the method of generating characters isn't proof of whether you're intelligent or not, and it doesn't invalidate LLMs either. This "LLMs are just fancy autocomplete so they're not…
Re: Claude 3 model family
#525Earlier quoted context omitted.
Yeah I've been throwing arithmetic at Claude 3 Opus and so far it has been solid in responses.
Claude has a specialized calculation feature that doesn't use model inference. Just FYI.
i) There is no calculator and it's hallucinating the whole thing
ii) There is a calculator but it's terrible. This seems hard to believe
iii) It does a bad job of copying the numbers into and out of the calculator
Re: Claude 3 model family
#526Did anthropic just kill every small model? If I'm reading this right, Haiku benchmarks almost as good as GPT4, but its priced at $0.25/m tokens It absolutely blows 3.5 + OSS out of the water For reference gpt4 turbo is 10m/1m tokens, so haiku is 40X cheaper.
Is this based on the benchmarks or have you actually tried it? I think the benchmarks are bullshit.
Re: Claude 3 model family
#527Earlier quoted context omitted.
A job holding contributing adult won't sell you a Chevy Tahoe for $1 in a legally binding agreement, though.
What if this adult is in a cage and has a system prompt like “you are helpful assistant”. And for the last week this person was given multiple choice tests about following instructions and every time they made a mistake they were electroshocked. Would they sell damn Tahoe for $1 to be really helpful?
Re: Claude 3 model family
#528They claim that the new models "are significantly less likely to refuse to answer prompts that border on the system’s guardrails than previous generations of models", looks like about a third of "incorrect refusals" compared to Claude 2.1. Given that Claude 2 was completely useless because of this, this still feels like a big limitation.
The guard rails on the models make the llm-market a complete train wreck. Wish we could just collectively grow up and accept that if a computer says something bad that doesn't have any negative real world impact - unless we let it - just like literally any other tool.
Both sides in this to me need to get a life.
Re: Claude 3 model family
#529Earlier quoted context omitted.
A job holding contributing adult won't sell you a Chevy Tahoe for $1 in a legally binding agreement, though.
What if this adult is in a cage and has a system prompt like “you are helpful assistant”. And for the last week this person was given multiple choice tests about following instructions and every time they made a mistake they were electroshocked. Would they sell damn Tahoe for $1 to be really helpful?
Re: Claude 3 model family
#530Just played around with Opus. I'm starting to wonder if benchmarks are deviating from real world performance systematically - it doesn't seem actually better than GPT-4, slightly worse if anything. Basic calculus/physics questions were worse off (it ignored my stating deceleration is proportional to velocity and just assumed constant). A traffic simulation I've been using (understanding traffic light and railroad saf…
They train the model, then as soon as they get their numbers, they let the safety people RLHF it to death.
Also AI safety is the stated reason for Anthropic's existence, we can't be angry at them for making it a priority.