Live data from Hacker News

Claude 3 model family

anthropic.com

201–210 of 723 posts

Re: Claude 3 model family

#201

Earlier quoted context omitted.

Arbitrary region locking : for example supported in Algeria and not in the neighboring Tunisia ... both are in North Africa

There's nothing arbitrary about it and both being located in North Africa means nothing. Tunisia has somewhat strict personal data protection laws and Algeria doesn't. That's the difference.

I know both countries, and in Algeria the Law No. 18-07, effective since August 10, 2023, establishes personal data protection requirements with severe penalties. The text is somewhat more strict than Tunisia.

Re: Claude 3 model family

#202

Earlier quoted context omitted.

Sure, OpenAI's next model would be expected to regain the lead, just due to their head start, but this level of catch-up from Anthropic is extremely impressive. Bear in mind that GPT-3 was published ("Language Models are Few-Shot Learners") in 2020, and Anthropic were only founded after that in 2021. So, with OpenAI having three generations under their belt, Anthropic came from nothing (at least in terms of models -…

What this really says to me is the indefensibility of any current advances. There’s really cool stuff going on right now, but anyone can do it. Not to say anyone can push the limits of research, but once the cat’s out of the bag, anyone with a few $B and dozen engineers can replicate a model that’s indistinguishably good from best in class to most users.

Yes, it seems that AI in form of LLMs is just an idea whose time has come. We now have the compute, the data, and the architecture (transformer) to do it.

As far as different groups leapfrogging each other for supremacy in various benchmarks, there might be a bit of a "4 minute mile" effect here too - once you know that something is possible then you can focus on replicating/exceeding it without having to worry are you hitting up against some hard limit.

I think the transformer still doesn't get the credit due for enabling this LLM-as-AI revolution. We've had the compute and data for a while, but this breakthough - shared via a public paper - was what has enabled it and made it essentially a level playing field for anyone with the few $B etc the approach requires.

I've never seen any claim by any of the transformer paper ("attention is all you need") authors that they understood/anticipated the true power of this model they created (esp. when applied at scale), which as the title suggests was basically regarded an incremental advance over other seq2seq approaches of the time. It seems like one of history's great accidental discoveries. I believe there is something very specific about the key-value matching "attention" mechanism of the transformer (perhaps roughly equivalent to some similar process used in our cortex?) that gives it it's power.

Re: Claude 3 model family

#204

Just signed up for Claude Pro to try out the Opus model. Decided to throw a complex query at it, combining an image with an involved question about SDXL fine tuning and asking it to do some math comparing the cost of using an RTX 6000 Ada vs an H100. It made a lot of mistakes. I provided it with a screenshot of Runpod's pricing for their GPUs, and it misread the pricing on an RTX 6000 ADA as $0.114 instead of $1.14.…

How many uses do you get per day of Opus with the pro subscription?

Re: Claude 3 model family

#205

I don't put a lot of stock on evals. many of the models claiming gpt-4 like benchmark scores feel a lot worse for any of my use-cases. Anyone got any sample output? Claude isn't available in EU yet, else i'd try it myself. :(

[deleted]

Re: Claude 3 model family

#206

What is the probability that newer models are just overfitting various benchmarks? A lot of these newer models seem to underperform GPT-4 in most of my daily queries, but I'm obviously swimming in the world of anecdata.

High. The only benchmark I look at is LMSys Chatbot Arena. Lets see how it perform on that

https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...

Re: Claude 3 model family

#207
I think to truly compete on the user side of things, Anthropic needs to develop mobile apps to use their models. I use the ChatGPT app on iOS (which is buggy as hell, by the way) for at least half the interactions I do. I won't sign up for any premium AI service that I can't use on the go or when my computer dies.

Re: Claude 3 model family

#208

Earlier quoted context omitted.

That's the good thing about intelligence: We have no fucking clue how to define it, so the goalpost just keeps moving.

In both directions. There are a set of people who are convinced that dolphins, octopi and dogs have intelligence, but GPT et al don't. I'm in the camp that says GPT4 has it. It's not a superhuman level of general intelligence, far from it, but it is a general intelligence that's doing more than regurgitation and rules-following.

How's a GPT not rules-following?

Re: Claude 3 model family

#209

I've tried all the top models. GPT4 beats everything I've tried, including Gemini 1.5- until today. I use GPT4 daily on a variety of things. Claude 3 Opus (been using temperature 0.7) is cleaning up. I'm very impressed.

Do you have specific examples?

Otherwise your comment is not quite useful or interesting to most readers as there is no data.

Re: Claude 3 model family

#210
post #41

Earlier quoted context omitted.

On the other hand, programmers are very expensive. At some level of accuracy and consistency (human order-of-magnitude?), the pricing of the service should start approaching the pricing of the human alternative. And first glance at numbers, LLMs are still way underpriced relative to humans.

NVidia's execs think so. It would be an ironic thing that it was open source that killed the programmer; as how would they train it otherwise? As a scientist, should I continue to support open access journals, just so I can be trained away? Slightly tongue in check, but not really.

> As a scientist, should I continue to support open access journals, just so I can be trained away?

If science was reproducible form articles posted in open access journals, we wouldn’t have half the problems we have with advancing research now.

Slightly tongue in check, but not really.

Post reply on HN