Live data from Hacker News

The Battle over Books3

wired.com

31–40 of 135 posts

Re: The Battle over Books3

#31
post #24

Earlier quoted context omitted.

Is all the data on the internet from 2010 to 2025 that much more valuable than all the data on the internet from 2010 to 2022? Better data yes, but more data up to the present day? Do we really need that to keep improving AI? GPT is already capable of incredible generalized language understanding and I'd wager we've long since hit diminishing returns from raw internet data. RLHF, fine tuning, and better (and more dat…

Language shifts, so NLP models need to understand those. For example, compare the use of "gay" before ~1980 and after. Some words change in spelling, like "to-morrow" present in works around 1800 changing to "tomorrow" in current usage. Some words are also coined, like "woke" or "enshitification". If your data or models don't account for those then it can make mistakes. For example, if a model is only trained on mode…

I always imagined a great source of training would be subtitles of live news or other live programming.

Would help to stay up to date on current events, and up to date on its general understanding of the world.

With the risk being adopting the bias of news media.

Re: The Battle over Books3

#32

From a preservation angle, how big is Books3? Is it easy enough for mortals to mirror for that unlikely future where it might be possible to self-train reasonably good models from scratch if provisioned with data?

books3.tar.gz itself is ~37gb compressed. Often really the entire "The Pile" dataset (composed of both the mostly compressed archives, along with a ~450gb compressed jsonl compilation of the data) is being discussed. That's around 825gb.

Out of curiosity why is jsonl so popular in the ML space?

Re: The Battle over Books3

#33

Any time I see the phrase "democratize access," my spidey-sense starts tingling. It's almost never used to describe an action that's an unadulterated good for society. It's USUALLY used to describe something sketchy at best, or even outright evil, with the justification that only "the bad guys" have access currently, and everything would be better if EVERYONE had access. Look, I get that unethical corporations using…

That's how I see it as well. To democratize art, music etc. now means to remove the skill component with the usage of all the work done so far. No one is actually prevented from pursuing those things and if you don't want your art to become training data you're ruining democratization and are somehow against the will of the people.

Re: The Battle over Books3

#34
What worries me most is that this is likely to only increase the gap between large corporate creator and small independent creators.

Large AI companies probably are not too worried about making deals with corporate content creators. Having access to content from a trigger-happy creator is only going to increase their advantage over competitors, after all. And if those creators were instead to try to introduce legislation, the AI companies would risk losing access to content from small creators without the means to sue too.

We seem to be moving into a world where corporate content cannot in any way be reused, remixed, or even archived. You cannot even own a copy - it is only accessible for a monthly fee and can disappear at any time. Meanwhile, anything created by independent creators is fair game to steal and rip off. Copyright was intended to promote and protect human creativity, but instead we got a rent-seeking mechanism used to stifle original creation.

Re: The Battle over Books3

#35
post #8

I wonder how this affects things like the Oxford English Dictionary, wiktionary, and other corpus-based linguistics, which rely on sample sentences of the word usage in order to determine the context. Because language is an evolving thing, it is almost certain that they have referenced sentences from copyrighted sources. E.g. I'm willing to bet that they have the sentence where Cory Doctorow introduces the term "ensh…

oed.com has 0 results for "enshittification" yet.

Re: The Battle over Books3

#36

Any time I see the phrase "democratize access," my spidey-sense starts tingling. It's almost never used to describe an action that's an unadulterated good for society. It's USUALLY used to describe something sketchy at best, or even outright evil, with the justification that only "the bad guys" have access currently, and everything would be better if EVERYONE had access. Look, I get that unethical corporations using…

Replace democratize with lower the barriers to entry. Yes it removes initial skill but typically doesn't remove the skill cap.

Re: The Battle over Books3

#37
post #20
post #9

Earlier quoted context omitted.

An LLM can be trained to find relevant knowledge online. It doesn’t have to be trained on the entirety of all existing knowledge.

Good luck when everything is paywalled

Given the already huge cost of training, and the evident lack of concern the LLM folks seem to have for copyright, why wouldn't the AI groups purchase subs to scrape the paywalled content?

The would possibly need to apply some effort to appear human, but that should only throttle the rate, not stop their scraping all together.

Re: The Battle over Books3

#38
post #4
post #2

What's next for models like GPT now that a lot of sites will outright block CCBot, GPTBot and others? How big of an impact is this going to have on the LLM itself? Isn't OpenAI in a bit of a pickle in regards to this? The problem with my question is the following: Content gets syndicated anyway, so if DigitalOcean blocks GPTBot (which it does), pretty much every single one of those tutorials will be syphoned off to o…

Laws. And the lawsuits that arise when such theft is discovered.

Who is going to enforce those laws?

Big publishers are more than happy to settle with AI companies - they just their slice of the pie after all. But who is going to protect, say, your Hacker News comments? Are you going to sue the AI company? Is YC going to sue? Are you going to sue YC for not banning crawlers in their robots.txt?

Re: The Battle over Books3

#39
> “It is only fair that you compensate us for using our writings, without which AI would be banal and extremely limited,” the letter states.

These models are going to come out so fast from one-off contracts. Its not the line that some creatives think it is. Its a 2 year delay at best.

Re: The Battle over Books3

#40
post #34

What worries me most is that this is likely to only increase the gap between large corporate creator and small independent creators. Large AI companies probably are not too worried about making deals with corporate content creators. Having access to content from a trigger-happy creator is only going to increase their advantage over competitors, after all. And if those creators were instead to try to introduce legisla…

How shocking! A monopoly right granted by the government to exclude others from use of an idea (patents) and creative works (intellectual property) was intended to help the little guy, but ended up helping the big boys extract rents instead, while exploiting the little guy because they sold their rights to survive and get access to a platform?

https://janefriedman.com/i-would-rather-see-my-books-pirated...

It’s almost as if governments (even capitalist ones) work with industry and concentration of power perpetuates this kind of consolidation further. They keep us distracted so we don’t have enough collective willpower to get together and demand reform, or even better — create our own alternative open ecosystems.

I write about many other examples of government-industry distracting us here: https://magarshak.com/blog/?p=362

Post reply on HN