Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

31–40 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#31
post #24
post #12

> Web pages crawled with the GPTBot user agent may potentially be used to improve future models > To disallow GPTBot to access your site you can add the GPTBot to your site’s robots.txt Too late - they already grabbed content from my personal website.

If they implemented this properly, they should be retroactively filtering all their content that is no longer allowed in the robots.txt, or carries the #NoAI tag.

My understanding is that it's not easy to untrain a model of data already fed to it.

Regarding noai tags - is this respected or just wishful?

    

Re: GPTBot – OpenAI’s Web Crawler

#32
post #27
post #21

Earlier quoted context omitted.

If I learn something from your StackOverflow answers, do you expect me to share a percentage of my future salary with you?

Can you share your knowledge with millions at once? If so, then pay.

So there's an infinite pyramid of "who learned what from who", and payment flows upwards along the hierarchy, all the way back to people who are long dead, and then down to their descendants who presumably inherited their "knowledge rights"?

You can't be serious. Thank god our world doesn't work like that.

Re: GPTBot – OpenAI’s Web Crawler

#33
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Every response so far is no i.e. the hobby website doesn't merit any compensation.

A contrarian take to support the original commenter is that if the site owner had ads, i probably got him or her some increment in site visits and helped in some small way with monetization, site ranking and boosted his / her public persona, credibility.

When GPT bot visits, none of that happens. Much worse - people who might have visited the hobby site and contributed to traffic and ad revenue will now start getting their answers from the OpenAI chatbot and never visit this hobby site.

That's exploitation and I think that's what most of the responses on this thread miss.

Re: GPTBot – OpenAI’s Web Crawler

#35
post #16
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…

The distinction is scale at which OpenAI can make profit off of your work. Now this might sound trivial, but it's scale of fraud possible has been the biggest argument against online elections.

Re: GPTBot – OpenAI’s Web Crawler

#38
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

scooba diving

This is a bit pedantic but the term is "scuba diving". Scuba is an acronym that's short for "self contained underwater breathing apparatus". It doesn't work if you don't spell it right.

Re: GPTBot – OpenAI’s Web Crawler

#39
Yet another bot that completely ignores the "429 Too Many Requests" response header and happily continues hammering your tiny little side project [1] to death. Luckily, I already block the IP address they're using as it has been used for (other?) malicious bots before.

[1] In my case, it relies on third-party APIs that are heavily rate limited. Any bot ignoring rate limitation measures will effectively (D)DOS my service.

Re: GPTBot – OpenAI’s Web Crawler

#40
post #32
post #27

Earlier quoted context omitted.

Can you share your knowledge with millions at once? If so, then pay.

So there's an infinite pyramid of "who learned what from who", and payment flows upwards along the hierarchy, all the way back to people who are long dead, and then down to their descendants who presumably inherited their "knowledge rights"? You can't be serious. Thank god our world doesn't work like that.

It's not about who learned what from whom, it's about superstar economy. If you serve all customers and leave nothing for the rest, it will be a problem.

What do you think why writers and actors have included AI in the reasons of their strike?

Post reply on HN