Live data from Hacker News

Claude 3 model family

anthropic.com

441–450 of 723 posts

Re: Claude 3 model family

#441
post #364

Opus (and all prior Claude models) still do not correctly answer the Sally question, unfortunately: Q: Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? A: Okay, let's organize the information we have: Sally is a girl. Sally has 3 brothers. Each of Sally's brothers has 2 sisters. Now, let's think about this step by step: We know that Sally is a girl, so she is a sister to he…

It seems like it is getting tripped up on grammar. Do these models not deterministically preparse text input into a logical notation?

No, they're a "next character" predictor - like a really fancy version of the auto-complete on your phone - and when you feed it in a bunch of characters (eg. a prompt), you're basically pre-selecting a chunk of the prediction. So to get multiple characters out, you literally loop through this process one character at a time.

I think this is a perfect example of why these things are confusing for people. People assume there's some level of "intelligence" in them, but they're just extremely advanced "forecasting" tools.

That said, newer models get some smarts where they can output "hidden" python code which will get run, and the result will get injecting into the response (eg. for graphs, math, web lookups, etc).

Re: Claude 3 model family

#442
The HumanEval benchmark scores are confusing to me.

Why does Haiku (the lowest cost model) have a higher HumanEval score than Sonnet (the middle cost model)? I'd expect that would be flipped. It gives me the impression that there was leakage of the eval into the training data.

Re: Claude 3 model family

#443
post #364

Opus (and all prior Claude models) still do not correctly answer the Sally question, unfortunately: Q: Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? A: Okay, let's organize the information we have: Sally is a girl. Sally has 3 brothers. Each of Sally's brothers has 2 sisters. Now, let's think about this step by step: We know that Sally is a girl, so she is a sister to he…

Yeah, cause these are the kinds of very advanced things we'll use these models for in the wild. /s

It's strange that these tests are frequent. Why would people think this is a good use of this model or even a good proxy for other more sophisticated "soft" tasks?

Like to me, a better test is one that tests for memorization of long-tailed information that's scarce on the internet. Reasoning tests like this are so stupid they could be programmed, or you could hook up tools to these LLMs to process them.

Much more interesting use cases for these models exist in the "soft" areas than 'hard', 'digital', 'exact', 'simple' reasoning.

I'd take an analogical over a logical model any day. Write a program for Sally.

Re: Claude 3 model family

#444
post #364

Opus (and all prior Claude models) still do not correctly answer the Sally question, unfortunately: Q: Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? A: Okay, let's organize the information we have: Sally is a girl. Sally has 3 brothers. Each of Sally's brothers has 2 sisters. Now, let's think about this step by step: We know that Sally is a girl, so she is a sister to he…

Seems stochastic? This is what I see from Opus which is correct: https://claude.ai/share/f5dcbf13-237f-4110-bb39-bccb8d396c2b Did you perhaps run this on Sonnet?

[deleted]

Re: Claude 3 model family

#445

Earlier quoted context omitted.

I cant wait until this is the true disruptor in the economy: " Take this $1,000 and maximise my returns and invest it where appropriate. Goal is to make this $1,000 100X " And just let your r/wallStreetBets BOT run rampant with it...

That will only work for the first few people who try it.

They will allow access to Ultimate version to X people only for just $YB/m charge.

Re: Claude 3 model family

#446

Earlier quoted context omitted.

Yes, it seems that AI in form of LLMs is just an idea whose time has come. We now have the compute, the data, and the architecture (transformer) to do it. As far as different groups leapfrogging each other for supremacy in various benchmarks, there might be a bit of a "4 minute mile" effect here too - once you know that something is possible then you can focus on replicating/exceeding it without having to worry are y…

> We now have the compute, the data, and the architecture (transformer) to do it. It's really not the model, it's the data and scaling. Otherwise the success of different architectures like Mamba would be hard to justify. Conversely, humans getting training on the same topics achieve very similar results, even though brains are very different at low level, not even the same number of neurons, not to mention different…

I don’t think we are at a plateau. We may have fed a large amount of text into these models, but when you add up all other kinds of media, images, videos, sound, 3D models, there’s a castle more rich dataset about the world. Sora showed that these models can learn a lot about physics and cause and effect just from video feeds. Once this is all combined together into multimodal mega models then we may be closer to the plateau.

Re: Claude 3 model family

#447
post #364

Opus (and all prior Claude models) still do not correctly answer the Sally question, unfortunately: Q: Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? A: Okay, let's organize the information we have: Sally is a girl. Sally has 3 brothers. Each of Sally's brothers has 2 sisters. Now, let's think about this step by step: We know that Sally is a girl, so she is a sister to he…

Groq's Mixtral 8x7b nails this one though. https://groq.com/ Sally has 1 sister. This may seem counterintuitive at first, but let's reason through it: We know that Sally has 3 brothers, and she is one of the sisters. Then we are told that each brother has 2 sisters. Since Sally's brothers share the same parents as Sally, they share the same sisters. Therefore, Sally's 3 brothers have only 1 additional sister besides…

If you change the names and numbers a bit, e.g. "Jake (a guy) has 6 sisters. Each sister has 3 brothers. How many brothers does Jake have?" it fails completely. Mixtral is not that good, it's just contaminated with this specific prompt.

In the same fashion lots of Mistral 7B fine tunes can solve the plate-on-banana prompt but most larger models can't, for the same reason.

https://arxiv.org/abs/2309.08632

Re: Claude 3 model family

#448

Earlier quoted context omitted.

Sure, but it would also be an IA much smarter than the ones we have now, because you cannot replace a human being with the current technology. You can augment one, making her perform the job of two or more humans before for some tasks, but you cannot replace them all, because the current tech cannot reasonably be used without supervision.

a lot of jobs are being replaced by AI already... comms/copywriting/customer service/off shored contract technicals roles especially.

No they aren't. Some jobs are being scaled down because of the increased productivity of other people with AI, but none of the jobs you listed are within reach of autonomous AI work with today's technology (as illustrated by the AirCanada hilarious case).

Re: Claude 3 model family

#449
post #430

This is my highly advanced test image for vision understanding. Only GPT-4 gets it right some of the time - even Gemini Ultra fails consistently. Can someone who has access try it out with Opus? Just upload the image and say "explain the joke." https://i.imgur.com/H3oc2ZC.png

This is what I got on the Anthropic console, using Opus with temp=0: > The image shows a cute brown and white bunny rabbit sitting next to a small white shoe or slipper. The text below the image says "He lost one of his white shoes during playtime, if you see it please let me know" followed by a laughing emoji. > The joke is that the shoe does not actually belong to the bunny, as rabbits do not wear shoes. The captio…

Thanks. This is about on par with what Gemini Ultra responds, whereas GPT-4 responds better (if oddly phrased in this run):

> The bunny has fur on its hind feet that resembles a pair of white shoes. However, one of the front paws also has a patch of white fur, which creates the appearance that the bunny has three "white shoes" with one "shoe" missing — hence the circle around the paw without white fur. The humor lies in the fact that the bunny naturally has this fur pattern that whimsically resembles shoes, and the caption plays into this illusion by suggesting that the bunny has misplaced one of its "shoes".

Re: Claude 3 model family

#450

"leading the frontier of general intelligence." Llms are an illusion of general intelligence. What is different about these models that leads to such a claim? Marketing hype?

Turing might disagree with you that it is an _illusion_.
Post reply on HN