Live data from Hacker News

Will we run out of ML data? Evidence from projecting dataset size trends (2022)

epochai.org

61–70 of 126 posts

Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)

#61
post #42

This is dumb. An individual human takes in more data than modern LLMs do. https://open.substack.com/pub/echoesofid/p/why-llms-struggle...

Blind kids don't though, and they still end up being smarter than GPT-4

Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)

#62

Brute force approaches always hit some wall. ML will be no different. In the decades to come it us quite likely that algorithms will develop in directions orthogonal to current approaches. The idea that you improve performance by throwing gazillions of data into gargantuan models might be even come to be seen as laughable. Keep in mind (pun) that the only real intelligence here is us, and we are pretty good at figuri…

We won't hit the wall. Somewhat counterintuitively, scaling datasets is the lazy and economical approach. If you have the compute already, might as well dig an OOM more text tokens. But there are other sources of data, and slightly different ways to utilize it. Multimodality, in very large training runs, will almost inevitably increase sample efficiency (for obvious reasons of context richness), synthetic data is alr…

But we can devise generally applicable rules of reasoning from first principles. It's called logic. I am pretty sure the next step is to properly combine machine learning and logic properly.

Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)

#63
post #23

Earlier quoted context omitted.

> All thanks to "democratizing" ip by stealing data. That horse is already dead. Large models can learn everything, there's nothing that can be done to stop them from learning. It's too easy for them to do it. We can't hold any meaningful IP when models can generate 100 variations only different enough to pass the test. IP is dead. But on its corpse there will grow a new world of applications. We all got new skills,…

Sure. In before people used to say that currency is dead because crypto currency has replaced money already, and already people are using them, and already [insert marketing statement]. I see they now moved on to ai.

>currency is dead because crypto currency has replaced money already

Literally no one said that, you are being ridiculous

Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)

#64

On this note, the data-set available if you start collecting today is tainted with experimental AI content. Not the biggest issue right now but as time goes on this problem will get worse and we'll be basing our simulations of intelligence on the output of our simulations of intelligence, a brave new abstraction.

I see this take a lot and I think it's quite wrong, not fully, but at least missing a couple big points.

I think they already don't blindly feed it just all the garbage raw data they can find, but prefer high quality, well-prepared sources.

And aside from spam, we're not just blindly posting AI content either. We're putting in meaningful prompts, rejecting answers we don't like, and editing answers we do.

Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)

#65

I wonder if the better question is not how we get more training data but: If we're running out of training data with hallucinations and performance remaining so inadequate (per OpenAI's whitepaper) is an autoregressive transformer the right architecture? Perhaps ongoing work in finetuning will take these models to the next level but ignoring the LLM hype it really does seem like things have plateaued for a while now…

I don't think that the hallucinations have anything to do with the architecture, rather they come from optimizing a cost function where saying "I don't know" is as bad as being wrong. I do not think that RLHF as currently understood can fix this, since the reward model would struggle to distinguish fact from fiction.

Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)

#66
post #44

Earlier quoted context omitted.

Uncompressed 2x 8k by 8k 24bpp 24FPS video. Comes at about 500GB per hour.

It’s always interesting to think about the exact technological analog for our biological sensors, but I believe that our vision would be way less than that in terms of raw data. We have a super high-res area at the center of vision (the fovea), but the rest is extremely low resolution (but with high movement and light sensitivity). I think we could reasonably say that if an optical nerve has 1mm neurons on average, a…

31mb/s = 14GB/hr (bits to bytes). 81TB per year, assuming 16 hours awake per day. Fits snuggly on a large SSD ;)

Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)

#68

Earlier quoted context omitted.

We won't hit the wall. Somewhat counterintuitively, scaling datasets is the lazy and economical approach. If you have the compute already, might as well dig an OOM more text tokens. But there are other sources of data, and slightly different ways to utilize it. Multimodality, in very large training runs, will almost inevitably increase sample efficiency (for obvious reasons of context richness), synthetic data is alr…

But we can devise generally applicable rules of reasoning from first principles. It's called logic. I am pretty sure the next step is to properly combine machine learning and logic properly.

Seems unlikely, that never worked in the past. And humans don't actually use logic (especially formal logic) to come up with anything. They just use it to justify what they came up with.

Not even mathematicians think in terms of logic when trying to solve problems.

Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)

#69

On this note, the data-set available if you start collecting today is tainted with experimental AI content. Not the biggest issue right now but as time goes on this problem will get worse and we'll be basing our simulations of intelligence on the output of our simulations of intelligence, a brave new abstraction.

Text on the Internet isn't just tainted by the output of recent relatively smart models.

We had computers spit out text (especially spam) for ages now. You'd have to filter those out, too, if tainting actually was a problem.

Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)

#70

On this note, the data-set available if you start collecting today is tainted with experimental AI content. Not the biggest issue right now but as time goes on this problem will get worse and we'll be basing our simulations of intelligence on the output of our simulations of intelligence, a brave new abstraction.

I see this take a lot and I think it's quite wrong, not fully, but at least missing a couple big points. I think they already don't blindly feed it just all the garbage raw data they can find, but prefer high quality, well-prepared sources. And aside from spam, we're not just blindly posting AI content either. We're putting in meaningful prompts, rejecting answers we don't like, and editing answers we do.

Um, really? You think your average 'growth hacker' who is using ChatGPT to exponentially increase the amount of SEO junk they can churn out is checking each answer before they press publish?

Purity, accuracy and relevance of data collected from the internet is going to a very hard problem.

Post reply on HN