Look at that jump in grade school math. From 55 % with GPT 3.5 to 95 % for both Claude 3 and GPT 4.
Yeah I've been throwing arithmetic at Claude 3 Opus and so far it has been solid in responses.
Claude 3 model family
161–170 of 723 posts
Re: Claude 3 model family
#162I don't put a lot of stock on evals. many of the models claiming gpt-4 like benchmark scores feel a lot worse for any of my use-cases. Anyone got any sample output? Claude isn't available in EU yet, else i'd try it myself. :(
Re: Claude 3 model family
#163Earlier quoted context omitted.
Yeah the output pricing I think is really interesting, 150% more expensive input tokens 250% more expensive output tokens, I wonder what's behind that? That suggests the inference time is more expensive then the memory needed to load it in the first place I guess?
Either something like that or just because the model's output is basically the best you can get and they utilize their market position. Probably that and what you mentioned.
Re: Claude 3 model family
#164I use GPT4 daily on a variety of things.
Claude 3 Opus (been using temperature 0.7) is cleaning up. I'm very impressed.
Re: Claude 3 model family
#165Re: Claude 3 model family
#166What's up with the weird list of the supported countries? It isn't available in most European countries (except for Ukraine and UK) but on the other hand lot of African counties are listed... https://www.anthropic.com/claude-ai-locations
Re: Claude 3 model family
#167Surpassing GPT4 is huge for any model, very impressive to pull off. But then again...GPT4 is a year old and OpenAI has not yet revealed their next-gen model.
MMLU is pretty much the only stat on there that matters, as it correlates to multitask reasoning ability. Here, they outpace GPT-4 by a smidge, but even that is impressive because I don’t think anyone else’s has to date.
It's genuinely outperforming GPT4 in my manual tests.
Re: Claude 3 model family
#168Earlier quoted context omitted.
Sure, OpenAI's next model would be expected to regain the lead, just due to their head start, but this level of catch-up from Anthropic is extremely impressive. Bear in mind that GPT-3 was published ("Language Models are Few-Shot Learners") in 2020, and Anthropic were only founded after that in 2021. So, with OpenAI having three generations under their belt, Anthropic came from nothing (at least in terms of models -…
What this really says to me is the indefensibility of any current advances. There’s really cool stuff going on right now, but anyone can do it. Not to say anyone can push the limits of research, but once the cat’s out of the bag, anyone with a few $B and dozen engineers can replicate a model that’s indistinguishably good from best in class to most users.
Re: Claude 3 model family
#169Just added Claude 3 to Chat at https://double.bot if anyone wants to try it for coding. Free for now and will push Claude 3 for autocomplete later this afternoon. From my early tests this seems like the first API alternative to GPT4. Huge!
Re: Claude 3 model family
#170Earlier quoted context omitted.
So far gpt is the only one able to answer to variations of these prompts https://www.lesswrong.com/posts/EHbJ69JDs4suovpLw/testing-pa... it might be trained on these but still you can create variations and get decent responses Most other model fail on basic stuff like the python creator on stack overflow question, they identify Guido as the python creator, so the knowledge is there, but they don't make the connection…
>>So far gpt is the only one able to answer to variations of these prompts You're saying that when Mistral Large launched last week you tested it on (among other things) explaining jokes?