Live data from Hacker News

Claude 3 model family

anthropic.com

161–170 of 723 posts

Re: Claude 3 model family

#162

I don't put a lot of stock on evals. many of the models claiming gpt-4 like benchmark scores feel a lot worse for any of my use-cases. Anyone got any sample output? Claude isn't available in EU yet, else i'd try it myself. :(

I've also seen the opposite, where tiny little 7B models get real close to GPT4 quality results on really specifically use cases. If you're trying to scale just that use case it's significantly cheaper, and also faster to just scale up inference with that specialty model. An example of this is using an LLM to extract medical details from a record.

Re: Claude 3 model family

#163
post #52

Earlier quoted context omitted.

Yeah the output pricing I think is really interesting, 150% more expensive input tokens 250% more expensive output tokens, I wonder what's behind that? That suggests the inference time is more expensive then the memory needed to load it in the first place I guess?

Either something like that or just because the model's output is basically the best you can get and they utilize their market position. Probably that and what you mentioned.

This. Price is set by value delivered and what the market will pay for whatever capacity they have; it’s not a cost + X% market.

Re: Claude 3 model family

#164
I've tried all the top models. GPT4 beats everything I've tried, including Gemini 1.5- until today.

I use GPT4 daily on a variety of things.

Claude 3 Opus (been using temperature 0.7) is cleaning up. I'm very impressed.

Re: Claude 3 model family

#166
post #99

What's up with the weird list of the supported countries? It isn't available in most European countries (except for Ukraine and UK) but on the other hand lot of African counties are listed... https://www.anthropic.com/claude-ai-locations

Arbitrary region locking : for example supported in Algeria and not in the neighboring Tunisia ... both are in North Africa

Re: Claude 3 model family

#167

Surpassing GPT4 is huge for any model, very impressive to pull off. But then again...GPT4 is a year old and OpenAI has not yet revealed their next-gen model.

MMLU is pretty much the only stat on there that matters, as it correlates to multitask reasoning ability. Here, they outpace GPT-4 by a smidge, but even that is impressive because I don’t think anyone else’s has to date.

I still don't trust benchmarks, but they've come a long way.

It's genuinely outperforming GPT4 in my manual tests.

Re: Claude 3 model family

#168

Earlier quoted context omitted.

Sure, OpenAI's next model would be expected to regain the lead, just due to their head start, but this level of catch-up from Anthropic is extremely impressive. Bear in mind that GPT-3 was published ("Language Models are Few-Shot Learners") in 2020, and Anthropic were only founded after that in 2021. So, with OpenAI having three generations under their belt, Anthropic came from nothing (at least in terms of models -…

What this really says to me is the indefensibility of any current advances. There’s really cool stuff going on right now, but anyone can do it. Not to say anyone can push the limits of research, but once the cat’s out of the bag, anyone with a few $B and dozen engineers can replicate a model that’s indistinguishably good from best in class to most users.

Barrier to entry with "few $B" is pretty high. Especially since the scaling laws indicate that it's only getting more expensive. And even if you manage to raise $Bs, you still need to be clever on how to deploy it (talent, compute, data) ...

Re: Claude 3 model family

#169

Just added Claude 3 to Chat at https://double.bot if anyone wants to try it for coding. Free for now and will push Claude 3 for autocomplete later this afternoon. From my early tests this seems like the first API alternative to GPT4. Huge!

Emacs implementation when? ;)

Re: Claude 3 model family

#170

Earlier quoted context omitted.

So far gpt is the only one able to answer to variations of these prompts https://www.lesswrong.com/posts/EHbJ69JDs4suovpLw/testing-pa... it might be trained on these but still you can create variations and get decent responses Most other model fail on basic stuff like the python creator on stack overflow question, they identify Guido as the python creator, so the knowledge is there, but they don't make the connection…

>>So far gpt is the only one able to answer to variations of these prompts You're saying that when Mistral Large launched last week you tested it on (among other things) explaining jokes?

Sorry I did what? When?
Post reply on HN