Live data from Hacker News

Large Enough

mistral.ai

61–70 of 512 posts

Re: Large Enough

#61
post #10

These companies full of brilliant engineers are throwing millions of dollars in training costs to produce SOTA models that are... "on par with GPT-4o and Claude Opus"? And then the next 2.23% bump will cost another XX million? It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; that an ARC prize level breakthrough is needed to move the needle any farther at this point.

The thing I don't understand is why everyone is throwing money at LLMs for language, when there are much simpler use cases which are more useful? For example, has anyone ever attempted image -> html/css model? Seems like it be great if I can draw something on a piece of paper and have it generate a website view for me.

> For example, has anyone ever attempted image -> html/css model?

There are already companies selling services where they generate entire frontend applications from vague natural language inputs.

https://vercel.com/blog/announcing-v0-generative-ui

Re: Large Enough

#62
post #39

Earlier quoted context omitted.

All 3 models you ranked cannot get "how many r's are in strawberry?" correct. They all claim 2 r's unless you press them. With all the training data I'm surprised none of them fixed this yet.

When using a prompt that involves thinking first, all three get it correct. "Count how many rs are in the word strawberry. First, list each letter and indicate whether it's an r and tally as you go, and then give a count at the end." Llama 405b: correct Mistral Large 2: correct Claude 3.5 Sonnet: correct

It’s not impressive that one has to go to that length though.

Re: Large Enough

#63
post #10

These companies full of brilliant engineers are throwing millions of dollars in training costs to produce SOTA models that are... "on par with GPT-4o and Claude Opus"? And then the next 2.23% bump will cost another XX million? It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; that an ARC prize level breakthrough is needed to move the needle any farther at this point.

> It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; I think you're just seeing the "make it work" stage of the combo "first make it work, then make it fast". Time to market is critical, as you can attest by the fact you framed the situation as "on par with GPT-4o and Claude Opus". You're seeing huge investments because being the first to get a working model stands to b…

ChatGPT is like Google now. It is the default. Even if Claude becomes as good as ChatGPT or even slightly better it won't make me switch. It has to be like a lot better. Way better.

It feels like ChatGPT won the time to market war already.

Re: Large Enough

#64

Earlier quoted context omitted.

GPT-4 was probably as good as Claude Sonnet 3.5 at its outset, but OpenAI ran it into the ground with whatever they’re doing to save on inference costs, otherwise scale, align it, or add dumb product features.

Indeed, it used to output all the code I needed but now it only outputs a draft of the code with prompts telling me to fill in the rest. If I wanted to fill in the rest, I wouldn't have asked you now, would've I?

It's doing something different for me. It seems almost desperate to generate vast chunks of boilerplate code that are only tangentially related to the question.

That's my perception, anyway.

Re: Large Enough

#66
post #55
post #3

This race for the top model is getting wild. Everyone is claiming to one-up each with every version. My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away. I'm not really sure how to even test/use Mistral or Llama for everyday use though.

Claude is pretty great, but it's lacking the speech recognition and TTS, isn't it?

Correct. IMO the official Claude app is pretty garbage. Sonnet 3.5 API + Open-WebUI is amazing though and supports STT+TTS as well as a ton of other great features.

Re: Large Enough

#67
post #10

These companies full of brilliant engineers are throwing millions of dollars in training costs to produce SOTA models that are... "on par with GPT-4o and Claude Opus"? And then the next 2.23% bump will cost another XX million? It seems increasingly apparent that we are reaching the limits of throwing more data at more GPUs; that an ARC prize level breakthrough is needed to move the needle any farther at this point.

The thing I don't understand is why everyone is throwing money at LLMs for language, when there are much simpler use cases which are more useful? For example, has anyone ever attempted image -> html/css model? Seems like it be great if I can draw something on a piece of paper and have it generate a website view for me.

Perhaps if we think of LLMs as search engines (Google, Bing etc) then there's more money to be made by being the top generic search engine than the top specialized one (code search, papers search etc)

Re: Large Enough

#68
post #33

When I see this "© 2024 [Company Name], All rights reserved", it's a tell that the company does not understand how hopelessly behind they are about to be.

Could you elaborate on this? Would love to understand what leads you to this conclusion.

E = T/A! [0]

A faster evolving approach to AI is coming out this year that will smoke anyone who still uses the term "license" in regards to ideas [1].

[0] https://breckyunits.com/eta.html [1] https://breckyunits.com/freedom.html

Re: Large Enough

#69
post #58

Earlier quoted context omitted.

Given how absolutely pitiful the proprietary advancements in AI have been, I would posit we have little to worry about.

OTOH the companies who are sharing their breakthroughs openly aren't yet making any money, so something has to give. Their research is currently being bankrolled by investors who assume there will be returns eventually, and eventually can only be kicked down the road for so long.

Eventually can be (and has been) bankrolled by Nvidia. They did a lot of ground-floor research on GANs and training optimization, which only makes sense to release as public research. Similarly, Meta and Google are both well-incentivized to share their research through Pytorch and Tensorflow respectively.

I really am not expecting Apple or Microsoft to discover AGI and ferret it away for profitability purposes. Strictly speaking, I don't think superhuman intelligence even exists in the domain of text generation.

Re: Large Enough

#70
post #3

This race for the top model is getting wild. Everyone is claiming to one-up each with every version. My experience (benchmarks aside) Claude 3.5 Sonnet absolutely blows everything away. I'm not really sure how to even test/use Mistral or Llama for everyday use though.

3.5 sonnet is the quality of the OG GPT-4, but mind blowingly fast. I need to cancel my chatgpt sub.

> mind blowingly fast

I would imagine this might change once enough users migrate to it.

Post reply on HN