Live data from Hacker News

Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

news.ycombinator.com

171–180 of 296 posts

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#171
post #45
post #14

Earlier quoted context omitted.

> Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. Failing to solve every problem does not mean a solution is a failure. From sunscreen to seatbelts, the world is full of great solutions that occasionally fail due to statistics and large numbers.

> Failing to solve every problem does not mean a solution is a failure. There is something to be said though to OP's point where it's actually better to do nothing than an AI.txt because it can give a false sense of security, which is obviously not what you want.

The point of an ai.txt is that it signals intention of the copyright holder.

Anytime a business is caught using that content, they can't claim that they used publicly available information, because the ai.txt specifically signalled to everyone in a clear and unambiguous manner that the copyright granted by viewing the page is witheld from ai training.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#173

If AI needs explicit information and context, surely it should focus on improving its context recognition rather than trying to fix that by inserting even more training data. Regardless, I do agree that something like a robots.txt for AI can be very useful. I'd like my website to be excluded from most AI projects and some kind of standardized way to communicate this preference would be nice, although I realize most A…

> focus on improving its context recognition rather than trying to fix that by inserting even more training data. That's how you improve its context recognition. You show it many contexts. > most AI projects don't exactly care about things like the wishes of authors, copyright, or ethical considerations Why is it 'ethical' that you get to add a bunch of restrictions to a pre-negotiated situation? You get copyright pr…

> There's a way to add restrictions - licensing - and you're looking to get the benefits of licensing, and to take away fair use right from other people, without paying the costs of doing so.

The way copyright laws work is that work is copyrighted (assuming the work is original enough, of course) by default. You don't get to use it unless you have a license. Now, of course, as an author, you can choose to add a license to your work (whether that's CC0 or GPL-3), but you don't have to.

You do have an implicit license to consume this content, but not to reproduce it. If you put all of those copies you've saved on some public other website, that's a copyright violation. Furthermore, access to privately-owned blog posts and websites is a privilege, not a right. You're not my boss, I don't have to write content for you.

The exact legal status of AI models trained on other people's unlicensed works and their output is still largely unknown. Legal professionals much more qualified than me have argued how AI models and generated work can either be completely fair use, with no need to apply any kind of copyright restriction, or how AI generated work can be classified as a derivative work, which means you need a license. There are two major lawsuits about this going on as far as I know and it'll take years for those to flesh out.

If it turns out that AI models and the works they produce are completely fair game, I suppose I'll need take down my content wherever I can in order not to be a free source of training data for big tech; public datasets and the internet archive will still have to respond to DMCA takedowns, after all. However, I'm not all that confident that what AI is doing is all that legally okay.

I have no problem with you saving and archiving anything you want to read. I also fully support the Internet Archive and its goal. I do have a problem with these multi billion dollar companies scouring the internet for their money maker, giving nothing in return.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#174

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

> The thing I somewhat struggle with is that after 20-30 years of calls for shorter copyright terms, lesser restrictions on content you access publicly, and what you can do with it, we are now in the situation where the arguments are quickly leaning the other way. "We" now want stricter copyright law when it comes to AI, but at the same time shorter copyright duration...

Are these mutually exclusive? If you couldn't make Avengers movie Thanos memes but all the 90s X-Men and Spiderman content was a free for all, I think a lot of people would take that trade off.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#175
robots.txt was a performance-hack. It never felt like a audience-filter. As sad as it might sound, hoping for filtering on publicly reachable content seems a bit naiv in my book. If you want your stuff not learnt by an AI, you better not publish it. Everything a human can read, an AI eventually will.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#176

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

There is no legal agreement to follow robots.txt, but it appears to have came up a few times (from the first search result for "court cases involving robots.txt"):

https://www.robotstxt.org/faq/legal.html

If an "ai.txt" were to exist, I hope it's a signal for opt-in rather than opt-out. Whereas "robots.txt" being an explicit signal for opt-out might be useful because people who build public websites generally want their websites to be discovered, it seemed unlikely that training unknown AI would be a use case that content creators had in mind, considering that most existing content predates current AI systems.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#177

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

With search engines and other crawlers, there wasn't easy ways to monetize "copyright theft" at scale. Google, which had the biggest share of eyeballs, was much more equitable in sharing revenue to content producers (who wanted to monetize). And Google was probably more just in taking action against copyright theft.

Individual high value IP was always much less accessible (not available as a webpage on the internet). Gen AI/LLMs with the internet scale data is too powerful and maybe easier to monetize.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#178
post #83

Earlier quoted context omitted.

> "We" now want stricter copyright law when it comes to AI, but at the same time shorter copyright duration... This gross generalization of other people's views on important issues is really offensive. My view is that the Copyright Act of 1976 had it about right when they established the duration of copyright. My view is that members of Congress were handsomely rewarded by a specific corporation to carve out special…

I think they nailed it with the original 1790 act. 14 years + 14 more is plenty.

My biggest critique of copyright is that is unnecessarily collapses financial reward & creative control. It also pegs both as starting at creation - which is not a particularly meaningful point for either problem.

IMO I would rather a structure that:

- Guarantees creators (and their descendants) some number of years of financial benefit / veto (30 seems fine!) - i.e. pay me what I want or you can't use this creative work.

- Separately grant creators the ability to veto "official" projects that use their creative output in their lifetimes.

IMO, it seems like there's a productive "middle ground" between total control and anything goes. After the 30 year benefit expired, you couldn't sue for damages - just costs & to stop use.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#179
post #83

Earlier quoted context omitted.

I think they nailed it with the original 1790 act. 14 years + 14 more is plenty.

Same. The very nature of information is that it yearns to be free. Information cannot be "owned." The point of copyright should be to grant temporary monopolies to encourage creation, not to confer ownership. Thomas Jefferson put it beautifully: If nature has made any one thing less susceptible than all others of exclusive property, it is the action of the thinking power called an idea, which an individual may exclus…

> The very nature of information is that it yearns to be free.

Information wants you to stop anthropomorphizing it.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#180
post #117

# cat > /var/www/.well-known/ai.txt Disallow: * ^D # systemctl restart apache2 Until then, I'm seriously considering prompt injection in my websites to disrupt the current generation of AI. Not sure if it would work. Please share with me ideas, links and further reading about adversarial anti-AI countermeasures. EDIT: I've made an Ask HN for this: https://news.ycombinator.com/item?id=35888849

I wouldn't want to be you when Roko's Basilisk emerges.
Post reply on HN