Live data from Hacker News

Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

news.ycombinator.com

201–210 of 296 posts

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#202
post #170

Earlier quoted context omitted.

but copyright is not for information or ideas, information and ideas cannot be copyrighted; it's for creative expression

and why should "creative expression" be owned?

Why should land be owned? None of us created the planet...

But we have selected an economic system that depends on ownership to drive exchange in a market, so... that's why.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#203
It's impossible.

The problem is that such ai.txt would be an unidimensional opinion based on what? On the way the site describes itself. So a self-referencing source.

But the AIs reading it, are precisely going to invariably be trained with different world views that will summarize and express opinions biased by these worldviews. It's even deeper as every worldview can't help but belong to one ideology or another.

So who is aligned with truth now?

The author? AI1? AI2? AI3?...AIN?

We're in such a mess.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#204
post #170

Earlier quoted context omitted.

and why should "creative expression" be owned?

Why should land be owned? None of us created the planet... But we have selected an economic system that depends on ownership to drive exchange in a market, so... that's why.

I'd argue that land is owned because it's a finite resource, and that without property ownership people would be in conflict with one another. "Creative expression" is not finite, in fact every human possesses it, it's also intangible, it's ideas, thoughts, ... , which I personally do not believe should be owned.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#206

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

I think you fundamentally misunderstood the OP's point. They're not trying to use their ai.txt as any sort of deterrent, legal or otherwise.

They are trying to use it as a form of extended metadata for training AIs. Essentially, "ah I see you're training using my website! Here's some extra info about it: [...]"

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#207
post #170

Earlier quoted context omitted.

but copyright is not for information or ideas, information and ideas cannot be copyrighted; it's for creative expression

and why should "creative expression" be owned?

Because creative expression can be exchanged for goods and services? Why should metal, wood, or special paper notes be owned? It's to represent work done and value to other people.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#208
post #83

Earlier quoted context omitted.

I think they nailed it with the original 1790 act. 14 years + 14 more is plenty.

Same. The very nature of information is that it yearns to be free. Information cannot be "owned." The point of copyright should be to grant temporary monopolies to encourage creation, not to confer ownership. Thomas Jefferson put it beautifully: If nature has made any one thing less susceptible than all others of exclusive property, it is the action of the thinking power called an idea, which an individual may exclus…

> very nature of information is that it yearns to be free. Information cannot be "owned."

The nature of information is to dissolve into entropy.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#210
At this point, all the good content has been sucked into LLM training sets. Other than a need to keep up with current events, there's no point in crawling more of the web to get training data.

There's a downside to dumping vast amounts of crap content into an LLM training set. The training method has no notion of data quality.

Post reply on HN