> One popular dataset, FineWeb, is about 11.25 trillion words long3, which, if you were reading at about 250 words per minute, would take you over 85 thousand years to read. It’s just not possible for any single human (or even a team of humans) to have read everything that an LLM has read during training. Do you have to read everything in a dataset with your own eyes to make sense of it? This would make any attempt t…
I mean, if you don't read it yourself, you're going to have to rely on _something/someone_ to filter/summarise the output, and at that point you might as well just accept that you'll never truly understand the entire thing?
I'll agree that we can do meaningful work (like reducing bias) without reading the entire dataset ourselves, but that doesn't reduce the fact that we cannot read everything that's going into these machines.