Large Enough
141–150 of 512 posts
Re: Large Enough
#142I'm building a ai coding assistant ( https://double.bot ) so I've tried pretty much all the frontier models. I added it this morning to play around with it and it's probably the worst model I've ever played with. Less coherent than 8B models. Worst case of benchmark hacking I've ever seen. example: https://x.com/WesleyYue/status/1816153964934750691
Re: Large Enough
#143These companies full of brilliant engineers are throwing millions of dollars in training costs to produce SOTA models that are... "on par with GPT-4o and Claude Opus"? And then the next 2.23% bump will cost another XX million? It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; that an ARC prize level breakthrough is needed to move the needle any farther at this point.
Re: Large Enough
#144Earlier quoted context omitted.
indeed. I pointed out in https://buttondown.email/ainews/archive/ainews-llama-31-the-... that the frontier model curve is currently going down 1 OoM every 4 months, meaning every model release has a very short half life[0]. however this progress is still worth it if we can deploy it to improve millions and eventually billions of people's lives. a commenter pointed out that the amoutn spent on Llama 3.1 was only like…
> however this progress is still worth it if we can deploy it to improve millions and eventually billions of people's lives Has there been any indication that we're improving the lives of millions of people?
Re: Large Enough
#145Earlier quoted context omitted.
Could you elaborate on this? Would love to understand what leads you to this conclusion.
E = T/A! [0] A faster evolving approach to AI is coming out this year that will smoke anyone who still uses the term "license" in regards to ideas [1]. [0] https://breckyunits.com/eta.html [1] https://breckyunits.com/freedom.html
Re: Large Enough
#146I'm building a ai coding assistant ( https://double.bot ) so I've tried pretty much all the frontier models. I added it this morning to play around with it and it's probably the worst model I've ever played with. Less coherent than 8B models. Worst case of benchmark hacking I've ever seen. example: https://x.com/WesleyYue/status/1816153964934750691
Re: Large Enough
#147What doe they mean by "single-node inference"? Do they mean inference done on a single machine?
Re: Large Enough
#148These companies full of brilliant engineers are throwing millions of dollars in training costs to produce SOTA models that are... "on par with GPT-4o and Claude Opus"? And then the next 2.23% bump will cost another XX million? It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; that an ARC prize level breakthrough is needed to move the needle any farther at this point.
Frankly I just don't understand the economics of training a foundation model. I'd rather own an airline. At least I can get a few years out of the capital investment of a plane.
Re: Large Enough
#149Re: Large Enough
#150This race for the top model is getting wild. Everyone is claiming to one-up each with every version. My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away. I'm not really sure how to even test/use Mistral or Llama for everyday use though.