This is dumb. An individual human takes in more data than modern LLMs do. https://open.substack.com/pub/echoesofid/p/why-llms-struggle...
Will we run out of ML data? Evidence from projecting dataset size trends (2022)
61–70 of 126 posts
Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)
#62Brute force approaches always hit some wall. ML will be no different. In the decades to come it us quite likely that algorithms will develop in directions orthogonal to current approaches. The idea that you improve performance by throwing gazillions of data into gargantuan models might be even come to be seen as laughable. Keep in mind (pun) that the only real intelligence here is us, and we are pretty good at figuri…
We won't hit the wall. Somewhat counterintuitively, scaling datasets is the lazy and economical approach. If you have the compute already, might as well dig an OOM more text tokens. But there are other sources of data, and slightly different ways to utilize it. Multimodality, in very large training runs, will almost inevitably increase sample efficiency (for obvious reasons of context richness), synthetic data is alr…
Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)
#63Earlier quoted context omitted.
> All thanks to "democratizing" ip by stealing data. That horse is already dead. Large models can learn everything, there's nothing that can be done to stop them from learning. It's too easy for them to do it. We can't hold any meaningful IP when models can generate 100 variations only different enough to pass the test. IP is dead. But on its corpse there will grow a new world of applications. We all got new skills,…
Sure. In before people used to say that currency is dead because crypto currency has replaced money already, and already people are using them, and already [insert marketing statement]. I see they now moved on to ai.
Literally no one said that, you are being ridiculous
Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)
#64On this note, the data-set available if you start collecting today is tainted with experimental AI content. Not the biggest issue right now but as time goes on this problem will get worse and we'll be basing our simulations of intelligence on the output of our simulations of intelligence, a brave new abstraction.
I think they already don't blindly feed it just all the garbage raw data they can find, but prefer high quality, well-prepared sources.
And aside from spam, we're not just blindly posting AI content either. We're putting in meaningful prompts, rejecting answers we don't like, and editing answers we do.
Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)
#65I wonder if the better question is not how we get more training data but: If we're running out of training data with hallucinations and performance remaining so inadequate (per OpenAI's whitepaper) is an autoregressive transformer the right architecture? Perhaps ongoing work in finetuning will take these models to the next level but ignoring the LLM hype it really does seem like things have plateaued for a while now…
Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)
#66Earlier quoted context omitted.
Uncompressed 2x 8k by 8k 24bpp 24FPS video. Comes at about 500GB per hour.
It’s always interesting to think about the exact technological analog for our biological sensors, but I believe that our vision would be way less than that in terms of raw data. We have a super high-res area at the center of vision (the fovea), but the rest is extremely low resolution (but with high movement and light sensitivity). I think we could reasonably say that if an optical nerve has 1mm neurons on average, a…
Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)
#67Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)
#68Earlier quoted context omitted.
We won't hit the wall. Somewhat counterintuitively, scaling datasets is the lazy and economical approach. If you have the compute already, might as well dig an OOM more text tokens. But there are other sources of data, and slightly different ways to utilize it. Multimodality, in very large training runs, will almost inevitably increase sample efficiency (for obvious reasons of context richness), synthetic data is alr…
But we can devise generally applicable rules of reasoning from first principles. It's called logic. I am pretty sure the next step is to properly combine machine learning and logic properly.
Not even mathematicians think in terms of logic when trying to solve problems.
Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)
#69On this note, the data-set available if you start collecting today is tainted with experimental AI content. Not the biggest issue right now but as time goes on this problem will get worse and we'll be basing our simulations of intelligence on the output of our simulations of intelligence, a brave new abstraction.
We had computers spit out text (especially spam) for ages now. You'd have to filter those out, too, if tainting actually was a problem.
Re: Will we run out of ML data? Evidence from projecting dataset size trends (2022)
#70On this note, the data-set available if you start collecting today is tainted with experimental AI content. Not the biggest issue right now but as time goes on this problem will get worse and we'll be basing our simulations of intelligence on the output of our simulations of intelligence, a brave new abstraction.
I see this take a lot and I think it's quite wrong, not fully, but at least missing a couple big points. I think they already don't blindly feed it just all the garbage raw data they can find, but prefer high quality, well-prepared sources. And aside from spam, we're not just blindly posting AI content either. We're putting in meaningful prompts, rejecting answers we don't like, and editing answers we do.
Purity, accuracy and relevance of data collected from the internet is going to a very hard problem.