Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

281–290 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#281
post #238

Earlier quoted context omitted.

If the cafe is a multi-billion corporation that can only exist because it can leech of content created by millions of other people without providing anything at all to them in return (and I'm not necessarily talking about financial compensation) then yeah.. maybe you should.

So Starbucks? Should they be sharing all their revenue with whoever invented all those Italian coffee drinks?

Starbucks business model is not entirely (or at all) reliant on the availability of new coffee drink recipes which can only be provided by third parties. So no, I wouldn't say so.

Re: GPTBot – OpenAI’s Web Crawler

#282
post #254

Earlier quoted context omitted.

Their papers say they were using Common Crawl for crawling. If you didn't want your pages in Common Crawl (eg. Twitter didn't) for use in many downstream analyses or uses beyond just OA, you could already have said so in your robots.txt.

Opt out != opt in. This reminds me of the beginning of the hitchhikers guide to the galaxy where Dent’s house is being demolished but the notice had been on display in a locked basement below city hall or something. He could have objected, technically!

I don't think that comparison is valid, and in fact, actually comparing them shows how reasonable it is: the HHGtG example is egregious because it is imposed silently, long after the fact, made deliberately invisible and hard to access, and discoverable only after the fact. All of those are false for robots.txt and Common Crawl. These are well-known, easy, old protocols which long predate most of the websites in question, which is completely disanalogous to the HHGtG example. Specifically: robots.txt precedes pretty much every website in existence. It's not some last-minute addition tacked on. Further, it is straightforward: you can deny scraping to everyone with a simple 'User-agent: * / Disallow: /' or nofollow headers (also 1 line in a web server like Apache or nginx) - hardly burdensome, and it rules out all projects, not just Common Crawl. Common Crawl is itself, incidentally, 15 years old, and long predates many of the websites it crawls, its crawler operates in the open with a clear user-agent and no shenanigans, and you can further look up what's in it because it's public. (This is how I know Twitter isn't in it: when people claimed GPT-3 was stealing answers from Twitter, I could just go check.) It is also well known, even many non-webmaster web users know about it because it governs what you'll see in search engines, what will be downloaded by some agents like wget by default, is covered early on in website materials, and so on.

Re: GPTBot – OpenAI’s Web Crawler

#283
post #16
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…

I don't think we really even need to dive that deep into the philosophical aspect of this. I think that it's fine to simply treat humans and machines differently, the same way we decided that animals cannot hold copyright for a work.

The reason copyright law exists in the first place is due to the difference of scale between copying books by hand and using a machine to do it, so I think "it's different because a machine is doing it" is a completely rational stance to take.

Re: GPTBot – OpenAI’s Web Crawler

#284
post #209

Earlier quoted context omitted.

That's of no benefit to me. Quite the contrary.

Why? If it helps people it should be good. Why bother posting something on the public web if not to help people. Sure a large org is receiving some ancillary benefit, but do you feel the same hostility for people working at [large corp] using what you worked on to help them at work? I honestly don't understand the hostility towards llms using public data

Don't understand or don't agree? Because it's really very simple to understand.

Generally people need some kind of incentive to produce content. This could be just the thought of somebody, an actual human, having consumed your content. Or a like, a comment exchange that further enriches the topic. Perhaps it leads to a new follower or even a new (online) friend. A job opportunity. Even a date. Or maybe just plain ad impressions to make your effort worthwhile.

The picture of content production was already bleak. Google gets to take it all for free and is the traffic controller deciding who gets the crumbs, and even then is also the sole advertiser. But at least they might throw you some traffic, leading to all the interactions I just mentioned.

OpenAI just steals your shit without permission, credit or payment and completely cuts of any direct human interaction with the original content or its maker.

How can you not "understand" the hostility? This is existential not just for the open web, also the closed web. Have you missed the developments at Twitter, StackOverflow, Reddit?

Re: GPTBot – OpenAI’s Web Crawler

#285
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Every response so far is no i.e. the hobby website doesn't merit any compensation. A contrarian take to support the original commenter is that if the site owner had ads, i probably got him or her some increment in site visits and helped in some small way with monetization, site ranking and boosted his / her public persona, credibility. When GPT bot visits, none of that happens. Much worse - people who might have visi…

The responses are nihilistic libertarian, as is typical here.

When a private company takes the sum of human knowledge without permission, attribution or payment and then monetizes it via the back door whilst cutting of any connection between the intended consumer and publisher, then we're dealing with a system I'd describe as criminal. It cannot be morally defended as "fair" in any major economical or political system.

The fact that they call it "Open" AI shows the level of trolling involved.

Re: GPTBot – OpenAI’s Web Crawler

#286

Earlier quoted context omitted.

I did this to Google for a while only to have my domains listed as malicious. I did not offer any malicious material, just different content for search engines was enough to flag my sites. They also did this to me when I gave google different IP addresses using a split DNS view. This was a while back so maybe they stopped this, I honestly don't know. Now I just give them and most bots a password prompt. Google and mo…

Dumb question but how would they know the content is different unless they're also crawling incognito and comparing the results?

Google have numerous robots that do not say Googlebot in the user-agent. They look just like Android cell phones. That is how they spot malicious sites or sites that are trying to game SEO or what-not. They are not within published CIDR blocks for Google and appear to just use wireless networks.

Re: GPTBot – OpenAI’s Web Crawler

#288
post #44
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Since that delta is clearly a transformative use of your photo -- as in, the output doesn't even remotely resemble the input -- you don't have any legal claim to it, no. I'm not sure what the plaintiffs arguing otherwise are smoking if they think they can argue it isn't transformative.

I don't think it's that simple. In order to have a claim to fair use, you would have to argue that the derivative work doesn't negatively affect the market for the original. When Google got sued for scanning copyrighted works for Google Books [1], they could claim fair use since they were only letting people see small excerpts from the books.

If you can train your bot on my blog post about scuba diving without my permission and then people can ask your bot for scuba diving advice instead of reading my blog, that doesn't seem very fair.

[1]: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

Re: GPTBot – OpenAI’s Web Crawler

#289

Earlier quoted context omitted.

> Depends who's reality On a trivial level this is correct as words mean what we collectively decide they mean. However I am making the point that a) the meaning has been changed and b) it has changed in a way that is deceptive and masks a useful fact about the world

> On a trivial level this is correct as words mean what we collectively decide they mean. Correct, and collectively we decided that reselling digital work without permission is indeed theft, just as we rightfully decided that digital goods for the most part are like physical goods. > a) the meaning has been changed It hasn't really, digital theft still has the same meaning as any form of theft. Some did try to change…

This debate predates modern AI and I've been having this debate for a lot longer than generative AI had been around. I think it's more likely that you really want to make a point about AI rather then you have deeply held views on intellectual property

Re: GPTBot – OpenAI’s Web Crawler

#290
post #44

Earlier quoted context omitted.

Since that delta is clearly a transformative use of your photo -- as in, the output doesn't even remotely resemble the input -- you don't have any legal claim to it, no. I'm not sure what the plaintiffs arguing otherwise are smoking if they think they can argue it isn't transformative.

I don't think it's that simple. In order to have a claim to fair use, you would have to argue that the derivative work doesn't negatively affect the market for the original. When Google got sued for scanning copyrighted works for Google Books [1], they could claim fair use since they were only letting people see small excerpts from the books. If you can train your bot on my blog post about scuba diving without my per…

> I don’t think it’s that simple. In order to have a claim to fair use, you would have to argue that the derivative work doesn’t negatively affect the market for the original.

No, you don’t.

That’s a factor weighing in favor of fair use, but the fair use factors are not defined in such a way that that is a necessary factor.

Post reply on HN