Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

291–300 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#291
post #16
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…

What if very rich people came to your small free-entry photo studio to look at your pictures, and - perhaps because they have very fast jets - also go to every other photo studio in the world to look at every other painter’s pictures? Knowing this, would you still let them in for free?

I believe no. Most people would make a distinction between “normal” and “rich”. They would give normal people free access, but the rich should pay for it.

It’s like a billionaire asking for a free hot dog. It’s like “come on, you can easily pay $100, which could even sponsor it for the next 100 people”.

Here it’s not the AI itself that’s exploiting you. It’s the rich people that make the AI that get even richer - partly thanks to your free work.

Re: GPTBot – OpenAI’s Web Crawler

#292
post #30

Earlier quoted context omitted.

The legal cases don't mean anything. The rule of law has all but disappeared from the corporate world. The idea that courts or regulators will be able to control AI is laughable. They are too corrupt, and they are way too slow.

The legal cases don't "mean anything" because AI training is /legal/, not because courts are "corrupt". If anything is transformative, an AI that doesn't memorize its input is.

Lossy compression of a 1MB original image into a 20kb compressed image doesn't make copyright go away

But that's essentially what LLMs are doing, lossy compression of the entire web

Re: GPTBot – OpenAI’s Web Crawler

#293
post #67

Is there any argument in favor of commercial websites allowing GPTBot to crawl them? It's not like Google where allowing crawling brings you traffic. In fact, it's pretty much the opposite.

Think about an AI as a personal assistant - for the whole world. Would you want the assistant to know about your business?

In most cases I do think so. It could mention you in a conversation, the analog to you appearing in Google Search results. And maybe even better, provide the necessary context to generate more real customers for your business. You don’t want traffic to your website, you want customers to your business. If you currently convert 10% of your traffic to customers, you’d be happy with 10% of the traffic of which you convert all to customers, because they are already converted before they even clicked your link.

Re: GPTBot – OpenAI’s Web Crawler

#294
post #172

I wonder — when GPTBot crawls my website, which has a number of translations performed using GPT, will it use all that data for training future models? That doesn't seem like a good idea, but I don't know how they could tell.

Models trained on the data from another model eventually leads to model collapse.

It’s like incest.

Re: GPTBot – OpenAI’s Web Crawler

#297

Earlier quoted context omitted.

Keeping you honest with incognito crawling is something they have to do anyway, to catch various tricks and scams - malware served up to users, etc.

So robots.txt is meaningless if they have to violate it to check for malicious content in blocked off pages anyways.

Well if you are blocking access to their crawler, I'd imagine they'd have no need to use an incognito crawler to check for malicious content. Why would they care if that content is not ending up in their index anyway?

Presumably, the incognito crawlers are only used on sites that have already granted the regular crawler access. That's content that ends up in their index which they want to vet.

Re: GPTBot – OpenAI’s Web Crawler

#298
post #17

Earlier quoted context omitted.

This doesn't appear consistent with other visitors to your website. If a cafe owner uses info on your site to improve their baking, should they also be required to share their revenue with you?

I see no point in engaging with this argument. ChatGPT is not a human. I, nor anyone else, should have to explain to you why that makes all the difference here.

If you don't want to engage in the argument, that's on you. I don't think ChatGPT not being a human makes any difference and I think the onus is on you to explain why it should.

Re: GPTBot – OpenAI’s Web Crawler

#300

Earlier quoted context omitted.

It’s win-win based on current usage. Even if OpenAI got shut down, I still benefited from using it. Many good things don’t last forever. If they go away that doesn’t invalidate the experiences you had.

I agree wrt/ experience, but I don't think it applies to this situation. Even if you had an experience that would end, their ownership of the data wouldn't, and that, among other things, make this very one-sided. I do want to stress something from your conclusion though. That people do better if they anticipate change, and can adapt to it.

Whether it's one-sided depends on what you think you've gained and lost. I publish code for free (open source) and I publish my writing for free (on my blog and as comments on various websites).

I don't expect compensation from anyone who uses them, whether it's public or private use, so I don't feel like I've lost anything. Sometimes people "pay it forward." If I actually get something back, that's a win.

There are web search engines and AI chatbots that might be very slightly better (unmeasurably so) due to having been trained on stuff I published over the years. Meanwhile I get a lot of benefit from using free stuff on the Internet. I think that's a one-sided deal in my favor.

(I also pay for GPT4 access. Whether it's worth $20 a month is more questionable, but it's fun to play with and so far I'm interested enough that I haven't cancelled.)

Post reply on HN