Live data from Hacker News

Compression is prediction

ngrok.com

281–290 of 324 posts

Re: Compression is prediction

#281

Earlier quoted context omitted.

They did not come up with the ideas themselves, so they got them somewhere. Follow the source and all the citations show up. It must be a modern thing where online blogging randos pretend they are all geniuses.

Better too assume they are just not aware. Technology is multi-layered cake of development. I have no doubt the only reason I know a lot of details is that I lived their development. When standing on the shoulders of giants it's hard to tell what is below them.

If they are not aware they are geniuses who rediscovered a bunch of shit in one blog post of effort. Better not to assume anything and let the writer tell you what's going on, if just an endnote

Re: Compression is prediction

#282
post #135
post #48

Earlier quoted context omitted.

I had a long ranting comment I deleted. I just don't like this trend of people presenting work in a way that makes you think some combo of 1) they discovered from scratch themselves 2) it's new 3) they didn't try to cite or acknowledge where they learned it/point to good sources 4) they don't really care about trying to teach something deeply, they want shiny stuff that makes them seem deep. This post references spec…

I think you'll find that 95% of all academic presentations are telling stories out of other peoples work.

Yes, and it's made pretty obvious with all the background section and references.. which btw I'm not advocating for such baggage in a blog post. Just something.

Re: Compression is prediction

#283
post #48

Earlier quoted context omitted.

I had a long ranting comment I deleted. I just don't like this trend of people presenting work in a way that makes you think some combo of 1) they discovered from scratch themselves 2) it's new 3) they didn't try to cite or acknowledge where they learned it/point to good sources 4) they don't really care about trying to teach something deeply, they want shiny stuff that makes them seem deep. This post references spec…

I don't think this is a fair critique. The author of the post uses standard terminology like entropy coding and arithmetic coding, and cited a paper "in 2023, Google DeepMind released a paper arguing that language modeling and compression are two views of the same thing" which discusses it further. This blog post is great. Well explained, and clearly took a lot of effort. I don't interpret it as them claiming to have…

I didn't say they claimed that. They left it rather vague, intentionally or not. I didnt mean to pick on just this post, I just happened to see a few such posts on the frontpage recently. I do think it was a good and obviously effortful blog post, overall.

Re: Compression is prediction

#284

Earlier quoted context omitted.

I completely agree, tremendous complexity can arise from simple mechanisms. Gleick's Chaos is a really good introduction to that. I was talking about the article though, and the mechanisms currently used for training and inference in LLMs. Those mechanisms are mathematically precise (unlike Chaotic attractors) and as the author points out, achieve the same function as compressors do in a strict bit pattern minimizati…

I do not understand this intuition that "true consciousness has to be random". The things that make me me are highly deterministic! > LLMs do not 'infer' token streams that haven't been trained in their training process While we're at it, this is simply untrue (in-context learning) unless you generalize "token streams" so radically that it could be readily analogized to humans as well.

> The things that make me me are highly deterministic!

Are they though? :-) There are some interesting papers in the tissue regeneration space which are working on building tissue (and organs) from stem cells for medical purposes (transplants, injury treatment, Etc.) and one of the things that comes out from that is that a set of stem cells make unique tissue every time in that it's compatible but the fine structure is always randomly different!

While the growth of brain matter is a minefield of ethical issues, at some point I suspect we're going to have to figure out how to do that to treat things like TBI and neurodegenerative diseases. In terms of understanding how randomness plays a part in your existence though cellular biology papers are a pretty good source.

Re: Compression is prediction

#285

Earlier quoted context omitted.

Citing 2023 makes it seem like this is newer than it is. Compression, prediction and intelligence have long been known to be deeply connected.

You expect every blog post to find the earliest relevant paper to cite, just so one could look at the year (without reading said paper - which would have made clear that the connection isn’t recent) to assess novelty? I don’t think that’s reasonable. It’s a blog post. If it was, say, a peer reviewed paper by Hinton or LeCunn that fails to cite Schmidhuber, that would be reasonable criticism in my opinion. (Spoiler: t…

This is why I hate online arguments. I didn't say earliest, but to give a pointer to the general era, and make it clear how standard the concepts are. They are part of undergrad, it's basic things. I'm not asking for a deep lit review. But modern exposition to anything related to Ai / ML / stats have extreme recency bias and young'uns are led to believe there was nothing before the transformer paper.

Re: Compression is prediction

#286
post #259

Consider: If you want to record the motion of the planets, naively you have large tables of positions. To compress that, you may smoothly interpolate sparse positions. To compress that, you encode the laws of gravity and simulate from a starting state. Compression is literally understanding.

Nice one. Yes. Compression constitutes necessary conditions for understanding. But don't forget recomposition as well. Add recursion to the mix and it results in recursive compression and recomposition - making understanding itself a self existing entity. This is as close to an understanding god you can rigorously get to.

This feels like an AI comment, but I feel compelled to respond.

Recursion is not needed when you regress system components to its foundamental representations, which can be derived without any recursion involved. In the thread example, I don't need to recursively determine how to compress how a planet orbits a mass. I only need to 'jump to the end' by define the rules of gravity and the mass/velocity of the bodies. What I'm driving at: I am suspicious of "Recursion is the Key" to better intelligence as it moves the goal post of understanding intelligence to 'mere' reflections without having to address other qualities of representations, qualia, novelty, etc.

Re: Compression is prediction

#287

Earlier quoted context omitted.

This shows that prediction algorithms (AI models) are also very good at compression, but compression algorithms (like the ones used in gzip) are not likewise very good at prediction. Which is evidence that compression is necessary but not sufficient for prediction.

LLMs are both the best compression and prediction algorithm for English text.

[deleted]

Re: Compression is prediction

#288

Earlier quoted context omitted.

I tried to reproduce those results, at least in terms of compression ratios, not speed. However I would say that testing on alice29, enwiki8, text8 data is kinda cheating. Alice in Wonderland and Wikipedia are very likely part of the training data of the LLM models used there. So I tried on HN comments from a few days ago, extracted from the text column of the public HN bigquery dataset. Using RWKV v7 0.1B instead of…

Oh wow, so it worked pretty well on data it hasn't seen. That expected but cool to reproduce. Have you seen this leaderboard of sorts[1], and this proposal to change hutter prize[2]? I think it's a really clever idea that you could measure an LLM's prediction abilities and language understanding by some sort of held-out compression metric because file sizes are very concrete. They are already beating shannon's number…

>I think it's a really clever idea that you could measure an LLM's prediction abilities and language understanding by some sort of held-out compression metric

This is the premise of https://huggingface.co/spaces/Jellyfish042/UncheatableEval

Re: Compression is prediction

#289
post #183

Earlier quoted context omitted.

> You expect every blog post to find the earliest relevant paper to cite This should be expected out of everyone. If you don't respect the reader enough to do this, why should we read your posts? I think papers should be retracted for not citing prior art, even if you weren't aware of it.

If the blog were about calculus, and stated that an elegant proof of the Fundamental Theorem of Calculus could be found in such and such undergraduate textbook, would you be upset that the citation wasn't to either Newton's or Leibniz' work?

Upset no.

But I would like it to be a violation of norms around citation. We should respect the reader and the truth.

Re: Compression is prediction

#290

Earlier quoted context omitted.

Your theory is that anybody who writes anything is obligated to make sure you can find any related information with one Google search? Again, to me that looks like wanting to be spoon fed. I have no idea why you think the world owes you endless 101-level discourse, but I hope you recognize you're setting yourself up for equally endless disappointment. If you take a little responsibility for your own education, you'll…

My theory is that anyone wishing to elucidate a complex technical concept is better served with contextual details. I think one method is superior to another. Anybody, anything, obligated, and any related information are hyperbolic terms I would have avoided. I understand with a weak position people often resort to hyperbole. I believe adding more terms to be googled is superior to fewer search terms. When working wi…

If you were trying to describe a preference, you didn't do a very good job. You stated it as a universal.

But taking you at your word, if there are things that don't meet your personal tastes, maybe move on to the next thing to read? Not everything has to be for everybody. There are plenty of audiences where people are capable of digging deeper when they want it. Personally, I prefer writers who don't spell everything out. The kind of writing you favor I generally find tedious. I'd much rather read something that shows a little faith in the audience.

Post reply on HN