Live data from Hacker News

Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

news.ycombinator.com

191–200 of 296 posts

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#191

Reading the title I thought you meant the opposite. Aka, an ai.txt file that disallow ai to train or use your data similar to robots.txt (but for cases when you still want to be crawled, just not extrapolated)

I've been (slowly) writing a new type of OSS license around this exact concept so it's easier to (legally) stop LLMs hoovering up IP [1] (under "derivative works not permitted").

[1] https://github.com/cheatcode/joystick/blob/development/LICEN...

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#192

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

"Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare." I like the idea of "ai.txt" but those who eat resources rarely listen to ToS. Frankly, I serve 503s to all identifiable bots, unless they are on my explicit allow list.

It might be a better idea to serve up a 418 ("I'm a tea pot") with a line line text file saying "I'm not an HTTP server". That solved a problem I had with bots making HTTP requests to my gopher server [1]. Serving up a 503 informs the bot that there's a server issue and it may try again later. A 418 informs the bot that it made an erroneous request and such an odd error code might get someone to look into it and stop.

[1] https://boston.conman.org/2019/09/30.2

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#193

Earlier quoted context omitted.

My phrasing was absolutely not meant to be read as myself speaking for all, apologies, I certainly don't want to offend. It has felt on HN and elsewhere that the prevailing attitude to copyright has been these two, somewhat contradictory, things. That's what I was trying to highlight with my phrasing of "we", which was also not meant to include myself but be a nod to the way a vocal group try to steer and dominate th…

Thank you! I think the average HN'er is frankly pretty ignorant about how copyright law works, the history around it, and the arguments for and against various reforms. In fairness it's an esoteric topic and most software developers depend in some way on copyrighted work for their income so that's not a huge surprise I guess. But it probably explains the contradiction you observed! The #1 issue with copyright today i…

There is contradiction because, in fact, the HN audience is more than one person and those people have different and conflicting views.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#194
post #186

Your HTML already has semantic meta elements like author and description you should be populating with info like that: https://developer.mozilla.org/en-US/docs/Learn/HTML/Introduc...

How do I add a semantic definition in an HTML tag to a JPEG, or MP4, or WAV, or any non HTML format? HTML tags fix HTML, not other formats.

JPEG has EXIF, MP3 has ID3 tags, MP4 has ilst, MKV has Tags, etc. We don't need xkcd/927 for these other formats that already have standard metadata mechanisms.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#195

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

> there is no "legal" agreement to follow those rules

Yes there certainly is[1]. The robots.txt clearly specifies authorized use and violating it exceeds that authorization. Now granted good luck getting the FBI to doorkick their friends at Google and other politically connected tech companies, but as the law is written crawlers need to honor the site owner's robots.txt.

[1] https://en.wikipedia.org/wiki/Computer_Fraud_and_Abuse_Act

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#196
post #186

Your HTML already has semantic meta elements like author and description you should be populating with info like that: https://developer.mozilla.org/en-US/docs/Learn/HTML/Introduc...

How do I add a semantic definition in an HTML tag to a JPEG, or MP4, or WAV, or any non HTML format? HTML tags fix HTML, not other formats.

If you're describing an object on the page, like an image or video, you want a label element linked to it by id and likely an aria-label attribute on the object. (screen readers and such will look for this in particular). For an image you want an alt text attribute as a description too.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#197
post #20
post #16

Earlier quoted context omitted.

> "We" now want stricter copyright law when it comes to AI, but at the same time shorter copyright duration... While I’m sure others than you share this opinion, I don’t think it’s as uniform as the more common “shorten/rationalize copyright terms and fair use” crowd “we.” I consider myself a knowledge worker and a pretty staunch proponent of floss and am perfectly fine with training AI on everything publicly availab…

Fair enough. But there should be some mechanism where people who don't want their works to contribute to AI training to be able to prevent that without having to resort to removing their works from the web.

I think people who don’t want their content contributing to AI shouldn’t have it on the public web.

There are many ways to restrict access. Use one of them. But if you respond to an anonymous http request with content then it shouldn’t matter if it’s a robot looking at it or a human (or a man or a woman or whatever).

I think this both for simplicity and that I foresee a future where human consciousness is simulated and basically an AI. I don’t want to have rules that biological humans can view and digital humans can’t.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#199

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

> "We" now want stricter copyright law when it comes to AI, but at the same time shorter copyright duration... This gross generalization of other people's views on important issues is really offensive. My view is that the Copyright Act of 1976 had it about right when they established the duration of copyright. My view is that members of Congress were handsomely rewarded by a specific corporation to carve out special…

> This gross generalization of other people's views on important issues is really offensive.

The "we" that has been calling for shorter terms is no more a gross generalization than the "we" that is calling for more protection against AI use of stuff.

The world outside of HN-and-similar has been much less anti-copyright than the world in here. More "neutral" seems to be dominant - we're not extending it anymore; we're not shrinking it either. And currently generally more panicked about AI taking away their jobs and rendering their skills and creativity useless.

The original post was a very fair summary of how there are now two ground-level movements competing that there weren't two years ago.

Post reply on HN