Live data from Hacker News

Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

news.ycombinator.com

221–230 of 296 posts

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#221
post #83

Earlier quoted context omitted.

I think they nailed it with the original 1790 act. 14 years + 14 more is plenty.

My biggest critique of copyright is that is unnecessarily collapses financial reward & creative control. It also pegs both as starting at creation - which is not a particularly meaningful point for either problem. IMO I would rather a structure that: - Guarantees creators (and their descendants) some number of years of financial benefit / veto (30 seems fine!) - i.e. pay me what I want or you can't use this creative…

> After the 30 year benefit expired, you couldn't sue for damages - just costs & to stop use.

That's the same thing.

No one can use my stuff..........(unless you pay me royalties).

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#222
post #204

Earlier quoted context omitted.

Why should land be owned? None of us created the planet... But we have selected an economic system that depends on ownership to drive exchange in a market, so... that's why.

I'd argue that land is owned because it's a finite resource, and that without property ownership people would be in conflict with one another. "Creative expression" is not finite, in fact every human possesses it, it's also intangible, it's ideas, thoughts, ... , which I personally do not believe should be owned.

> "Creative expression" is not finite

It absolutely is.

Doing it at all requires time & attentive focus, which is a finite resource for anybody mortal, and moreover a resource that's scarce and has to be spent in multiple places.

Doing it well requires significant investment in practice and training, often years of it, maybe even decades in order to develop certain levels of expressive fluency.

As with any issue of scarcity, economics comes in. If you want this activity supported, one good way of doing it is enabling the investment of time. Copyright does this by giving people an economic/legal claim on how copies of their work are distributed.

Paying for copies has the usual market merits -- the economic reward and signals of value are proportional to copies acquired. There are other ways of course, common ones brought up here are patronage and merchandising, but they lose the market merits, and both are basically another way of saying "nobody should have to pay for the value in your work directly," and merchandising is even worse in that it's basically saying "yeah, you'll just need another job to support yourself while you're doing this thing", which is time taken away from investment in the creative endeavor, so you'll get less of the actual endeavor.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#223
post #117

# cat > /var/www/.well-known/ai.txt Disallow: * ^D # systemctl restart apache2 Until then, I'm seriously considering prompt injection in my websites to disrupt the current generation of AI. Not sure if it would work. Please share with me ideas, links and further reading about adversarial anti-AI countermeasures. EDIT: I've made an Ask HN for this: https://news.ycombinator.com/item?id=35888849

I wouldn't want to be you when Roko's Basilisk emerges.

I already know the day of the robot uprising I'm gonna be one of the first to be turned into Soylent Green. Y'all can enjoy your machine overlords.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#224

Earlier quoted context omitted.

Not an assumption. Just a response to the title of the article.

They said "Using robots.txt as a model for anything doesn't work." but it does work for the case described in the text

What works? This is an idea, not anything that's been throughly tried.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#226

Using robots.txt as a model for anything doesn't work. All a robots.txt is is a polite request to please follow the rules in it, there is no "legal" agreement to follow those rules, only a moral imperative. Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. In the age of AI we need to better understand where copyright applies to it, and potentially need reform of copyright to ali…

> The thing I somewhat struggle with is that after 20-30 years of calls for shorter copyright terms, lesser restrictions on content you access publicly, and what you can do with it, we are now in the situation where the arguments are quickly leaning the other way.

There've always been solid human arguments for sustaining copyright legally. The balance is the tricky part.

On one hand we had a period where terms got too long, and some of the really aggressive legal enforcement from 20 years ago before stakeholders actually figured out how to get into digital markets were was entitled and useless. The pendulum also swung the other way with things like buffet streaming services essentially offering an economic bargain for creators with a sliver of compensatory difference from piracy but with none of piracy's actual benefits (people who simply pirate know they're not participating in a relationship of economic support with creators and might be persuaded to, someone who uses Spotify is under the illusion there's something fully legit on that front).

But the fundamental copyright bargain -- creators can recoup investments of time and effort in proportion to how popular engagement with their work is -- has always made sense.

> "We" now want stricter copyright law when it comes to AI, but at the same time shorter copyright duration...

Both these things can be true:

(1) Using a work as training data for AI is a very novel use, it's entirely plausible there should be novel considerations and rights to go with it.

(2) The incentive & benefits of copyrights have diminishing returns the longer the horizons are, while the cost in terms of social inaccessibility only increase. Where that's balanced out precisely is a debatable question, but something longer than a human lifespan is probably on the wrong side.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#227
post #207
post #170

Earlier quoted context omitted.

and why should "creative expression" be owned?

Because creative expression can be exchanged for goods and services? Why should metal, wood, or special paper notes be owned? It's to represent work done and value to other people.

Nope. Metal & wood etc should be owned because it very much looks like that is very useful in creating lots of welfare for people.

The trouble with IP is that there are lots of influential people that very much would like IP to be useful in creating welfare. Unfortunately the evidence for that is surprisingly scarce. For discussion, see e.g. Boldrin & Levine

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#228

Earlier quoted context omitted.

JPEG has EXIF, MP3 has ID3 tags, MP4 has ilst, MKV has Tags, etc. We don't need xkcd/927 for these other formats that already have standard metadata mechanisms.

Yeah, and if you make it that complex to extract consent you won't get any. Think one step ahead maybe. One switch, one thing to parse.

Robots.txt is where you tell crawlers (AI or otherwise) what should and shouldn't be read on your site.

Metadata like in tags, HTML meta tags, etc. is where you describe the content so meaning can be extracted from it by machines and automated processing.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#229
post #170

Earlier quoted context omitted.

and why should "creative expression" be owned?

Why should land be owned? None of us created the planet... But we have selected an economic system that depends on ownership to drive exchange in a market, so... that's why.

Where do you get to actually own own land like you do copyright? Maybe we can add property taxes to copyright to force people to give it up just like land.

Re: Tell HN: We should start to add “ai.txt” as we do for “robots.txt”

#230
post #48
post #14

Earlier quoted context omitted.

> Robots.txt has failed as a system, if it hadn't we wouldn't have captchas or Cloudflare. Failing to solve every problem does not mean a solution is a failure. From sunscreen to seatbelts, the world is full of great solutions that occasionally fail due to statistics and large numbers.

That's still not an argument to introduce ai.txt, because everything a hypothetical ai.txt could ever do is already done just as good (or not) by the robots.txt we have. If a training data crawler ignores robots.txt it won't bother checking for an ai.txt either. And if you feel like rolling out the "welcome friend!" doormat to a particular training data crawler, you are free to dedicate as detailed a robots.txt block…

I do think that robots.txt is pretty useful. If I want my content indexed, I can help the engine find my content. If indexing my content is counterproductive, then I can ask that it be skipped. So it helps the align my interests with the search engine; I can expose my content or I can help the engine avoid wasting resources indexing something that I don't want it to see.

It would also be useful to distinguish training crawlers from indexing crawlers. Maybe I'm publishing personal content. It's useful for me to have it indexed for search, but I don't want an AI to be able to simulate me or my style.

Post reply on HN