Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

171–180 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#171
post #30

Earlier quoted context omitted.

The legal cases don't mean anything. The rule of law has all but disappeared from the corporate world. The idea that courts or regulators will be able to control AI is laughable. They are too corrupt, and they are way too slow.

The legal cases don't "mean anything" because AI training is /legal/, not because courts are "corrupt". If anything is transformative, an AI that doesn't memorize its input is.

> If anything is transformative, an AI that doesn't memorize its input is.

I suspect the answer to the question "is it, though?" is one for the lawyers and lawmakers rather than for the software developers, and it may well vary wildly by jurisdiction.

Re: GPTBot – OpenAI’s Web Crawler

#172
I wonder — when GPTBot crawls my website, which has a number of translations performed using GPT, will it use all that data for training future models? That doesn't seem like a good idea, but I don't know how they could tell.

Re: GPTBot – OpenAI’s Web Crawler

#173
post #172

I wonder — when GPTBot crawls my website, which has a number of translations performed using GPT, will it use all that data for training future models? That doesn't seem like a good idea, but I don't know how they could tell.

Models trained on the data from another model eventually leads to model collapse.

Re: GPTBot – OpenAI’s Web Crawler

#174
post #102

Earlier quoted context omitted.

If you think copyright lawyers and the entertainment industry is going to let some AI upstarts launder their IP without a fight you aren't paying attention.

> AI upstarts You mean corporations that wield more power than most governments, and have revenues equivalent to the GDP of entire countries? If Universal or 20th Century Fox were to ever become a serious obstacle, Google and Microsoft are simply going to buy them. This isn't the early 2000s anymore. The power balance has shifted dramatically .

[deleted]

Re: GPTBot – OpenAI’s Web Crawler

#175
post #161

If I grab a copy of Adobe Photoshop (yeah, I know it runs in on a remote computer nowadays called 'the cloud') and I use it not to create the creative content its meant to be used for (manipulating cat pics, obviously) but to run it through IDA or Ghidra, or to study and use it to create a competitor (GIMP or make GIMP more like Photoshop) then even though I don't use it for its primary purpose; it is still copyright…

> Photoshop [...] runs in on a remote computer Does it? Last time I checked, “cloud” in “Creative Cloud” meant “now you have to pay a monthly subscription”. And reverse engineering Photoshop to make a competitor might be a legal practice, if done properly – for example, see the ReactOS project. Aside from that, I think your point still stands though.

Most of Photoshop works offline, but some of the newer 'AI' features run on Adobe servers and need an online connection (and account) to work.

Re: GPTBot – OpenAI’s Web Crawler

#176
post #16
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…

I know very few altruist humans. Whenever someone puts up some content online I believe there is always some motive from the author to benefit themselves even if it's subconscious. Perhaps through ad revenue or exposure from their blog/OSS project or just the dopamine of fake internet points from answering questions on forums. A human may particularly like your content and keep coming back to it or spread it with attribution.

But you don't get any of that from an LLM.

Re: GPTBot – OpenAI’s Web Crawler

#177
post #102

Earlier quoted context omitted.

> AI upstarts You mean corporations that wield more power than most governments, and have revenues equivalent to the GDP of entire countries? If Universal or 20th Century Fox were to ever become a serious obstacle, Google and Microsoft are simply going to buy them. This isn't the early 2000s anymore. The power balance has shifted dramatically .

FAANG already haven't bought or started competitors to the record labels they resell in their music stores. Don't see why they'll start now.

I just looked it up because I have no idea how big the music industry is, and…

US$26.2 billion globally in 2022 according to IFPI, and US$31.2 billion according to Statista.

Other than Netflix, I think FAANG just doesn't care that much about such a small market (the market being "actually producing it", given they're already part of the previous numbers for selling and streaming it).

And of course, both A's and the N of FAANG have their own commissioned TV/film content.

Re: GPTBot – OpenAI’s Web Crawler

#178
post #22
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

You don't have to publish anything on the internet. And when you do, you may limit the allowed audience to just the group of your friends etc. Why publish anything if you worry that someone may consume it?

I'm sure this technology is going to dissuade some people from publishing. Why bother if it is going to be regurgitated to everyone and their dog for $10 a month.

Re: GPTBot – OpenAI’s Web Crawler

#180
post #39

Yet another bot that completely ignores the "429 Too Many Requests" response header and happily continues hammering your tiny little side project [1] to death. Luckily, I already block the IP address they're using as it has been used for (other?) malicious bots before. [1] In my case, it relies on third-party APIs that are heavily rate limited. Any bot ignoring rate limitation measures will effectively (D)DOS my serv…

One option is to completely ban openai’s crawler ip addresses. They steal content without credit anyway - as most ai companies do - so there’s no benefit in allowing them access.

>so there’s no benefit in allowing them access.

Well, you're helping improving the model.

Post reply on HN