Live data from Hacker News

Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

news.ycombinator.com

131–140 of 296 posts

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#132
> It can be great if the website somehow ends up in a training dataset (who knows), and it can be super helpful for AI website crawlers, instead of using thousands of tokens to know what your website is about, they can do it with just a few hundred.

How do you differentiate an AI crawler from a normal crawler? Almost all of the LLMs are trained on commoncrawl, which the concept of LLMs didn't even exist when CC started. What about a crawler that creates a search database, but's context is fed into a LLM as context? Or a middleware that fetches data in real time?

Honestly that's a terrible idea. and robots.txt can cover the use cases. But is still pretty ineffective, because it's more just a set of suggestions than rules that must be followed.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#133
Attempts to muster and legitimize the ownership, squandering and sequestration of The Commons are growing rampant after the recent successes of generative AI. They are a tragic and misguided attempt to lesion, fragment and own the very consistency of The Collective Mind. Individuals and groups already have fairly absolute authority over their information property -- simply choose not to release it to The Commons. If you do not want people to see, sit or sleep on your couch, please keep it locked inside your home.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#134
post #83

Earlier quoted context omitted.

> "We" now want stricter copyright law when it comes to AI, but at the same time shorter copyright duration... This gross generalization of other people's views on important issues is really offensive. My view is that the Copyright Act of 1976 had it about right when they established the duration of copyright. My view is that members of Congress were handsomely rewarded by a specific corporation to carve out special…

I think they nailed it with the original 1790 act. 14 years + 14 more is plenty.

I would settle for 14 + 14 too :)

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#135
Why are we so defensive concerning human created content vs robot created content? Do we really need to feel frightened by some gpt?

Whilst the output of AI is astonishing by itself, is it really creating meaningful content en masse? I see myself relying more and more on human-curated content because typical commercialized use cases of AI generated stuff (product descriptions, corp blogs, SEO landing pages, etc.) all read like meaningless blabber, to me at least.

Whenever I see some cool techbro boasting how he created his "SEO factory" using ChatGPT, I can't help but think that the poor guy is shitting where he eats without even realizing it. Take Google with their Search and Ads; over the last decade they managed to bring down overall quality of web content that much, that I'm just completely fed up using it because by 99% chance I'll land on some meaningless SEO page.

From what I can perceive with things like HN, Mastodon, etc. it feels more like a rejuvenation of the human centric brand trusted Web. And by that I mean: Dear crawler, just use my content. Maybe you can do something good with it, maybe not. But chances are low, it's gonna replace me in any way but rather improve my content. It only leads to a downward spiral if we stick with the past of commercial thinking (more cheap content, more followers, more ads); if we'd instead switch to subscription models individuals won't get rich but we'd have a great ecosystem of ideas and content again.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#136

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

The point of robots.txt is to inform well behaved scrapers about how to behave. It is not designed nor intended to prevent bad actors.

Which is good design: don't pretend to solve problems you can't.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#137

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

It's not actually contradictory at all when you consider the root of the issue is about power asymmetry between individual creators and the corporations. Copyright terms were lobbied for, and primarily benefit the large corporations. They're symbolic of corporate overreach, that's why they're unpopular.

Meanwhile, now that the laws are inconvenient for them, tech companies are straight up ignoring labeling their training data to respect IP law. Labeling the data would be expensive, thereby eroding profits. The loss of usable data would also harm the efficacy of their models, and the time spent classifying the data will hamper their iteration time.

The ideas are only dissonant if you are looking at the trees (copyright term, DMCA, right to repair, etc.) and not the forest: which is a class struggle between a few thousand billionaires versus the rest of humanity.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#138
post #36

Earlier quoted context omitted.

Ok, fair point, I may be being a little hyperbolic. But my point is that it's not a system that we should copy for preventing the use of content in training AI. It would become a useless distraction. If you "violate" a robots.txt the server administrator can choose to block your bot (if they can fingerprint it) or IP (if its static). With an ai.txt there is no potential downside to violating it - unless we get new le…

> It's not a system that we should copy for preventing the use of content in training AI I don't see the OP saying anything about "ai.txt" being for that? They're advocating it as a way that AIs could use fewer tokens to understand what a site is about. (Which I also don't think is a good idea, since we already have lots of ways of including structured metadata in pages, but the main problem is not that crawlers woul…

Not only do we already have lots of ways of including structured metadata, but if you want to include directives about what should/shouldn't be scraped and by whom, we already have robots.txt.

In other words, there's no need to create an ai.txt when the robots.txt standard can just be extended.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#139

Earlier quoted context omitted.

> "We" now want stricter copyright law when it comes to AI, but at the same time shorter copyright duration... This gross generalization of other people's views on important issues is really offensive. My view is that the Copyright Act of 1976 had it about right when they established the duration of copyright. My view is that members of Congress were handsomely rewarded by a specific corporation to carve out special…

Why does Disney have an "ill-gotten" monopoly? The people who worked for the company created something. Why shouldn't they get to control how it's used. Do you feel like you should have control over what you create? Why not others?

Circular reasoning. If you assume your ideas are your own, and nobody else can benefit from them without your permission, then the point of your rhetorical questions follows. The reality is that IP laws are a grafting of property-like attributes onto something that absolutely isn't property.

Do I feel I should have control over what I create? I make hammers for a living. I sell them for $10. I don't expect any control over what people do with "my" hammers once I sell them. I don't even expect to stop my neighbor from buying one, teaching herself to build hammers, and then manufacturing and selling identical ones for $9. Do you?

(To anticipate the rest of this tired conversation, the temporary monopoly tradeoff ("securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries") is facially reasonable. But it's important to recognize that the "shouldn't" and "feel" in your questions are based on a very recent recharacterization of these temporary monopolies as "intellectual property," which is probably the most financially successful propaganda term ever devised. Start with "temporary monopoly" instead, and then the better rhetorical question for you to be asking is "when should Disney's temporary monopoly end?")

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#140

Earlier quoted context omitted.

> But my point is that it's not a system that we should copy for preventing the use of content in training AI The purpose OP is suggesting in the submission is the opposite, help AI crawlers to understand what the page/website is about without actually having to infer the purpose from the content itself.

Isn't that the entire point of the semantic web?

If only there was an HTML tag that let you provide a concise description of the page content. Perhaps something like
Post reply on HN