Live data from Hacker News

Big LLMs weights are a piece of history

antirez.com

151–160 of 222 posts

Re: Big LLMs weights are a piece of history

#151

Earlier quoted context omitted.

And the US ‘small’ LLMs will actually be slightly larger than the ‘large’ LLMs in the UK.

It's funny you say that, but when travelling abroad I wondered how Europeans and Japanese stay sufficiently hydrated.

Is this a thing about how restaurants in some European countries charge for water?

Re: Big LLMs weights are a piece of history

#152
post #56

> Scientific papers and processes that are lost forever as publishers fail, their websites shut down. I don't think the big scientific publishers (now, in our time) will ever fail, they are RICH!

That means nothing. Big companies fail all the time. There is no guarantee any of them will be here in 50 years, let alone 500.

Re: Big LLMs weights are a piece of history

#153

Earlier quoted context omitted.

Why not LLLM for large LLM’s and SLLM for small LLM’s, assuming there is no middle ground

What makes it a Small Large Language Model? Why jot just an SLM?

S and L cancel out, so it just an LM.

Re: Big LLMs weights are a piece of history

#154

Earlier quoted context omitted.

I'd prefer to see olive sizes get a renaissance. I was always amused by Super Colossal when following my mom around a store as a little kid. From a random web search, it seems the sizes above Large are: Extra Large, Jumbo, Extra Jumbo, Giant, Colossal, Super Colossal, Mammoth, Super Mammoth, Atlas.

And I'd love to see data compression terminology get an overhaul. Do we need big LLMs or just succinct data structures? Or maybe "compact" would be good enough? (Yeah LLMs are cool but why not just, you know, losslessly compress the actual data in a way that lets us query its content?)

Well the obvious answer is that LLMs are more then just pure search. They can synthesize novel information from their learned knowledge.

Re: Big LLMs weights are a piece of history

#155

Earlier quoted context omitted.

No. And every video game every made is available for download as well. If you even have to download it: they pride in making many of them playable in browser with just a click. Copyright issues aside (let's avoid that mess) I was referring to basic technical issues with the site. Design is atrocious, search doesn't work, you can click 50 captures of a site before you find one that actually loads, obvious data corrupt…

And yet, it's the best we currently have. I donate to them. We can come with demands of how it should be managed, but it should not prevent us from helping them.

If you poke around at what US government agencies are doing, and what European countries and non-profits are doing, or even do a deep dive into what your local library offers, you may find they no longer lead the pack.

They didn't even ask for donations until they accidentally set fire to their building annex. People offered to help (SF was apparently booming that year) and of course they promptly cranked out the necessary PHP to accept donations.

Now it's become part of the mythology. But throwing petty cash at a plane in a death spiral doesn't change gravity. They need to rehabilitate their reputation and partner with organizations who can help them achieve their mission over the long term. I personally think they need to focus on archival, legal long-term preservation and archival, before sticking their neck out any further. If this means no more Frogger in the browser, so be it.

I certainly don't begrudge anyone who donates, but asking for $17 on the same page as copyrighted game ROMs and glitchy scans of comic books isn't a long-term strategy.

Re: Big LLMs weights are a piece of history

#157

Earlier quoted context omitted.

Apple Intelligence has an LLM that runs locally on the iPhone (15 Pro and up). But the quality of Apple Intelligence shows us what happens when you use a tiny ultra-low-wattage LLM. There’s a whole subreddit dedicated to its notable fails: https://www.reddit.com/r/AppleIntelligenceFail/top/?t=all One example of this is “Sorry I was very drunk and went home and crashed straight into bed” being summarized by Apple Inte…

I think the real problem with LLMs is we have deterministic expectations of non-deterministic tools. We’ve been trained to expect that the computer is correct. Personally, I think the summaries of alerts is incredibly useful. But my expectation of accuracy for a 20 word summary of multiple 20-30 word summaries is tempered by the reality that’s there’s gonna be issues given the lack of context. The point of the summar…

Larger LLMs can summarize all of this quite well though.

Re: Big LLMs weights are a piece of history

#158

Earlier quoted context omitted.

It's funny you say that, but when travelling abroad I wondered how Europeans and Japanese stay sufficiently hydrated.

For healthy adults, thirst is a perfectly adequate guide to hydration needs. Historically normal patterns of drinking - e.g. water with meals and a few cups of tea or coffee in between - are perfectly sufficient unless you're doing hard physical labour or spending long periods of time outdoors in hot weather. The modern American preoccupation with constantly drinking water is a peculiar cultural phenomenon with no sc…

Diabetes causes dehydration

Re: Big LLMs weights are a piece of history

#159
post #138
post #87

Earlier quoted context omitted.

Doesn't the first L in LLM mean large already? It's like saying Automated ATM. Whoever wrote it barely knows what the acronym means. This whole article feels like written by someone who doesn't understand the subject matter at all

Yes, that's the point of the comment and the whole discussion here. LLMs are already Large so what should the prefix be? Big LLM is a strong contender. I'm also pretty sure the creator of redis is not "someone who doesn't understand the subject matter at all".

It's very common for experts on one subject to take a jab at another subject and depend on their reputation while their skillset doesn't translate at all.

Re: Big LLMs weights are a piece of history

#160
I would be curious to know if it would be possible to recunstruct approximate versions of popular common subsets of internet training data by using many different LLMs that may have happened to read the same info. Anyone knows pointers to math papers about such things?
Post reply on HN