Live data from Hacker News

Large Enough

mistral.ai

141–150 of 512 posts

Re: Large Enough

#142

I'm building a ai coding assistant ( https://double.bot ) so I've tried pretty much all the frontier models. I added it this morning to play around with it and it's probably the worst model I've ever played with. Less coherent than 8B models. Worst case of benchmark hacking I've ever seen. example: https://x.com/WesleyYue/status/1816153964934750691

to be fair that's quite a weird request (the initial one) – I feel a human would struggle to understand what you mean

Re: Large Enough

#143
post #10

These companies full of brilliant engineers are throwing millions of dollars in training costs to produce SOTA models that are... "on par with GPT-4o and Claude Opus"? And then the next 2.23% bump will cost another XX million? It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; that an ARC prize level breakthrough is needed to move the needle any farther at this point.

There is different directions AI have lots to improve: multi modal which branch into robotics, single modal like image, video, and sound generation and understanding. Also would check back when openAI releases 5

Re: Large Enough

#144
post #118
post #31

Earlier quoted context omitted.

indeed. I pointed out in https://buttondown.email/ainews/archive/ainews-llama-31-the-... that the frontier model curve is currently going down 1 OoM every 4 months, meaning every model release has a very short half life[0]. however this progress is still worth it if we can deploy it to improve millions and eventually billions of people's lives. a commenter pointed out that the amoutn spent on Llama 3.1 was only like…

> however this progress is still worth it if we can deploy it to improve millions and eventually billions of people's lives Has there been any indication that we're improving the lives of millions of people?

Yes, just like internet, power users have found use cases. It'll take education / habit for general users

Re: Large Enough

#145
post #68

Earlier quoted context omitted.

Could you elaborate on this? Would love to understand what leads you to this conclusion.

E = T/A! [0] A faster evolving approach to AI is coming out this year that will smoke anyone who still uses the term "license" in regards to ideas [1]. [0] https://breckyunits.com/eta.html [1] https://breckyunits.com/freedom.html

So it's made up?

Re: Large Enough

#146

I'm building a ai coding assistant ( https://double.bot ) so I've tried pretty much all the frontier models. I added it this morning to play around with it and it's probably the worst model I've ever played with. Less coherent than 8B models. Worst case of benchmark hacking I've ever seen. example: https://x.com/WesleyYue/status/1816153964934750691

What was the expected outcome for you? AFAIK, Python doesn't have a const dictionary. Were you wanting it to refactor into a dataclass?

Re: Large Enough

#148
post #10

These companies full of brilliant engineers are throwing millions of dollars in training costs to produce SOTA models that are... "on par with GPT-4o and Claude Opus"? And then the next 2.23% bump will cost another XX million? It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; that an ARC prize level breakthrough is needed to move the needle any farther at this point.

And with the increasing parameter size, the main winner will be Nvidia.

Frankly I just don't understand the economics of training a foundation model. I'd rather own an airline. At least I can get a few years out of the capital investment of a plane.

Re: Large Enough

#150
post #3

This race for the top model is getting wild. Everyone is claiming to one-up each with every version. My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away. I'm not really sure how to even test/use Mistral or Llama for everyday use though.

It’s these kind of praise that makes me wonder if they are all paid to give glowing reviews, this is not my experience with sonnet at all. It absolutely does not blow away gpt4o.
Post reply on HN