Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

161–170 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#161
If I grab a copy of Adobe Photoshop (yeah, I know it runs in on a remote computer nowadays called 'the cloud') and I use it not to create the creative content its meant to be used for (manipulating cat pics, obviously) but to run it through IDA or Ghidra, or to study and use it to create a competitor (GIMP or make GIMP more like Photoshop) then even though I don't use it for its primary purpose; it is still copyright infringement.

Same with this crawling by bots (Google, Bing, Meta, OpenAI; doesn't matter). Jurisprudence on Google News and Google Cache seems to show citing is OK, if done in moderation. Remember: just because you can access (download) something on the internet (WWW or otherwise) does not mean you're allowed to watch, use, save it. That argument was lost during the battles of copyright infringement in the years of 2000s.

OpenAI isn't even citing in moderation. Its making a derivative work without citing (hence obscuring) it does.

The bottom line is this: ML which doesn't cite sources should be regarded as hostile: a blackbox, and a copyright infringement paradise.

Re: GPTBot – OpenAI’s Web Crawler

#162
We'll se a change on this imho. LLMs are rationalising agents, not knowledge bases. We'll shift towards knowledge bases for LLMs that can and will attribute the source when returning answers. The LLM searches the knowledge base for you. I think the whole discussion is moot.

I'm working on something like that as well.

Re: GPTBot – OpenAI’s Web Crawler

#164
post #21
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

If I learn something from your StackOverflow answers, do you expect me to share a percentage of my future salary with you?

But are you human or not ? Because rights and laws that apply to humans do not necessarily apply to objects and vice versa. I don't expect a building permit from you when you stand on a piece of land. LLM's aren't legal entities in the formal sense, are they ?

Re: GPTBot – OpenAI’s Web Crawler

#165
post #31
post #24

Earlier quoted context omitted.

If they implemented this properly, they should be retroactively filtering all their content that is no longer allowed in the robots.txt, or carries the #NoAI tag.

My understanding is that it's not easy to untrain a model of data already fed to it. Regarding noai tags - is this respected or just wishful?

Untraining may be difficult, but will they only ever improve the current model? Never want to change its dimensions or parameters (I'm not too into the jargon) and train the fresh and improved version?

I'm not sure this first reasonably working chat bot is going to be the last version we ever need, and afaik this sort of thing is as hard to port as it is to untrain, the problem in both cases being that it's a big black box

Re: GPTBot – OpenAI’s Web Crawler

#166
Would blocking these bots from your site give bots that don't honor this a competitive advantage? Does that indirectly result in promoting not honoring robots.txt?

I'm considering whether to add it to my own site, but given that the future is already here and, while it's shitty to steal and regurgitate content without attribution at minimum, it's also not a big deal for my hobby site. It may serve my interests better to not include crawling restrictions for ClosedAI specifically

Re: GPTBot – OpenAI’s Web Crawler

#169
post #130

Earlier quoted context omitted.

I think a key idea is that with the amount of jurisdictions and number of courts the odds that a clean and sympathetic judge can be found approach one. I would argue that European jurisdictions are inherently less likely to be in pockets of American corporate interest and they are more likely to hear cases where fundamental human freedoms are at stake because both of these are existential threats to European independ…

The corporations that provide AI hold all the power because people (and businesses!) want to use their products. Let's say the French government decides that OpenAI must change something about their business practices if they want to continue operating in France. OpenAI says "nope", and blocks access to French users. Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countri…

> Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countries can. How long do you think it will take before a storm of corporate outrage forces the government to relent?

Bof, les alternatives à ChatGPT ne sont pas si mal.

And even if the open source alternatives were far behind rather than just a bit — all this talk about corporate moats and their absence may be blind to the strengths of OpenAI's offerings, but even so it can be replaced if it must — the storms of protest in France are normally by the people, not by the corporations.

Re: GPTBot – OpenAI’s Web Crawler

#170
post #161

If I grab a copy of Adobe Photoshop (yeah, I know it runs in on a remote computer nowadays called 'the cloud') and I use it not to create the creative content its meant to be used for (manipulating cat pics, obviously) but to run it through IDA or Ghidra, or to study and use it to create a competitor (GIMP or make GIMP more like Photoshop) then even though I don't use it for its primary purpose; it is still copyright…

> Photoshop [...] runs in on a remote computer

Does it? Last time I checked, “cloud” in “Creative Cloud” meant “now you have to pay a monthly subscription”.

And reverse engineering Photoshop to make a competitor might be a legal practice, if done properly – for example, see the ReactOS project.

Aside from that, I think your point still stands though.

Post reply on HN