Live data from Hacker News

Large Enough

mistral.ai

81–90 of 512 posts

Re: Large Enough

#81
post #58

Earlier quoted context omitted.

Given how absolutely pitiful the proprietary advancements in AI have been, I would posit we have little to worry about.

OTOH the companies who are sharing their breakthroughs openly aren't yet making any money, so something has to give. Their research is currently being bankrolled by investors who assume there will be returns eventually, and eventually can only be kicked down the road for so long.

Well, that's because the potential reward from picking the right horse is MASSIVE and the cost of potentially missing out is lifelong regret. Investors are driven by FOMO more than anything else. They know most of these will be duds but one of these duds could turn out to be life changing. So they will keep bankrolling as long as they have the money.

Re: Large Enough

#82
post #31
post #10

These companies full of brilliant engineers are throwing millions of dollars in training costs to produce SOTA models that are... "on par with GPT-4o and Claude Opus"? And then the next 2.23% bump will cost another XX million? It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; that an ARC prize level breakthrough is needed to move the needle any farther at this point.

indeed. I pointed out in https://buttondown.email/ainews/archive/ainews-llama-31-the-... that the frontier model curve is currently going down 1 OoM every 4 months, meaning every model release has a very short half life[0]. however this progress is still worth it if we can deploy it to improve millions and eventually billions of people's lives. a commenter pointed out that the amoutn spent on Llama 3.1 was only like…

Agreed on everything, but calling the marvel movies slop…I think that word has gone too far.

Re: Large Enough

#83
post #48

The question I (and I suspect most other HN readers) have is which model is best for coding? While I appreciate the advances in open weights models and all the competition from other companies, when it comes to my professional use I just want the best. Is that still GPT-4?

My personal experience says Claude 3.5 Sonnet.

The benchmarks agree as well.

Re: Large Enough

#84
post #5

Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/ I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history. I'd rank as: 1. Sonnet 3.5 2. Large 2 and Llama 405b (similar, no clear winner between the two) If you're using Claude, stick with it. My Claude wishlist: 1. Smarter (yes, it's the most intelligent, and y…

All 3 models you ranked cannot get "how many r's are in strawberry?" correct. They all claim 2 r's unless you press them. With all the training data I'm surprised none of them fixed this yet.

Lots of replies mention tokens as the root cause and I’m not well versed in this stuff at the low level but to me the answer is simple:

When this question is asked (from what the models trained on) the question is NOT “count the number of times r appears in the word strawberry” but instead (effectively) “I’ve written ‘strawbe’, now how many r’s are in strawberry again? Is it 1 or 2?”.

I think most humans would probably answer “there are 2” if we saw someone was writing and they asked that question, even without seeing what they have written down. Especially if someone said “does strawberry have 1 or 2 r’s in it?”. You could be a jerk and say “it actually has 3” or answer the question they are actually asking.

It’s an answer that is _technically_ incorrect but the answer people want in reality.

Re: Large Enough

#85
post #63

Earlier quoted context omitted.

> It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; I think you're just seeing the "make it work" stage of the combo "first make it work, then make it fast". Time to market is critical, as you can attest by the fact you framed the situation as "on par with GPT-4o and Claude Opus". You're seeing huge investments because being the first to get a working model stands to b…

ChatGPT is like Google now. It is the default. Even if Claude becomes as good as ChatGPT or even slightly better it won't make me switch. It has to be like a lot better. Way better. It feels like ChatGPT won the time to market war already.

But plenty people switched to Claude, esp. with Sonnet 3.5. Many of them in this very thread.

You may be right with the average person on the street, but I wonder how many have lost interest in LLM usage and cancelled their GPT plus sub.

Re: Large Enough

#86
post #78
post #34

Earlier quoted context omitted.

Meta just claimed the opposite in their Llama 3.1 paper. Look at the conclusion. They say that their experience indicates significant gains for the next iteration of models. The current crop of benchmarks might not reflect these gains, by the way.

I sell widgets. I promise the incalculable power of widgets has yet to be unleashed on the world, but it is tremendous and awesome and we should all be very afraid of widgets taking over the world because I can't see how they won't. Anyway here's the sales page. the widget subscription is so premium you won't even miss the subscription fee.

That is strong (and fun) point, but this is peer reviewable and has more open collaboration elements than purely selling widgets.

We should still be skeptical because often want to claim to be better or have unearned answers, but I don't think the motive to lie is quite as strong as a salesman's.

Re: Large Enough

#87
post #5

Links to chat with models that released this week: Large 2 - https://chat.mistral.ai/chat Llama 3.1 405b - https://www.llama2.ai/ I just tested Mistral Large 2 and Llama 3.1 405b on 5 prompts from my Claude history. I'd rank as: 1. Sonnet 3.5 2. Large 2 and Llama 405b (similar, no clear winner between the two) If you're using Claude, stick with it. My Claude wishlist: 1. Smarter (yes, it's the most intelligent, and y…

All 3 models you ranked cannot get "how many r's are in strawberry?" correct. They all claim 2 r's unless you press them. With all the training data I'm surprised none of them fixed this yet.

I wrote and published a paper at COLING 2022 on why LLMs in general won't solve this without either 1. radically increasing vocab size, 2. rethinking how tokenizers are done, or 3. forcing it with constraints:

https://aclanthology.org/2022.cai-1.2/

Re: Large Enough

#88
post #34

Earlier quoted context omitted.

Meta just claimed the opposite in their Llama 3.1 paper. Look at the conclusion. They say that their experience indicates significant gains for the next iteration of models. The current crop of benchmarks might not reflect these gains, by the way.

LLMs are reaching saturation on even some of the latest benchmarks and yet I am still a little disappointed by how they perform in practice. They are by no means bad, but I am now mostly interested in long context competency. We need benchmarks that force the LLM to complete multiple tasks simultaneously in one super long session.

I don't know anything about AI but there's one thing I want it to do for me. Program a full body exercise program long term based on the parameters I give it such as available equipment and past workout context goals. I haven't had good success with chatgpt but I assume what you're talking about is relevant to my goals.

Re: Large Enough

#89

Earlier quoted context omitted.

I think GPT5 will be the signal of whether or not we have hit a plateau. The space is still rapidly developing, and while large model gains are getting harder to pick apart, there have been enormous gains in the capabilities of light weight models.

> I think GPT5 will be the signal of whether or not we have hit a plateau. I think GPT5 will tell if OpenAI hit a plateau. Sam Altman has been quoted as claiming "GPT-3 had the intelligence of a toddler, GPT-4 was more similar to a smart high-schooler, and that the next generation will look to have PhD-level intelligence (in certain tasks)" Notice the high degree of upselling based on vague claims of performance, and…

>some degree of salesmanship

buddy every few weeks one of these bozos is telling us their product is literally going to eclipse humanity and we should all start fearing the inevitable great collapse.

It's like how no one owns a car anymore because of ai driving and I don't have to tell you about the great bank disaster of 2019, when we all had to accept that fiat currency is over.

You've got to be a particular kind of unfortunate to believe it when sam altman says literally anything.

Post reply on HN