Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

121–130 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#121
post #16
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…

Even if training a model turns out to be similar to human learning, I don't think it necessarily follows that it should be treated the same, legally or morally. There's nothing wrong with human laws or morals that enshrine human behavior, like the human way of learning, as special and distinct from machine learning.

Re: GPTBot – OpenAI’s Web Crawler

#122

What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…

Chatgpt 4 provides pretty good citations on request.

I don't think it's technically able to do that. It just tries to "guess" what the right source. It might get it right more often that not but that's not exactly what a citation is.

Re: GPTBot – OpenAI’s Web Crawler

#123
post #68

What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…

Just one example: As a software vendor, you probably want OpenAI to index your documentation, so questions about your software can be answered by ChatGPT. Not everybody who creates content is a "content creator" (when did this word get the specific meaning "people who earn money or reputation from creating content").

That’s a good point, I hadn’t thought of these cases.

On a second thought, I guess a lot of marketing content would also love to be crawled by anything that crawls…

Re: GPTBot – OpenAI’s Web Crawler

#124
post #56

Earlier quoted context omitted.

Interesting point though I'd go with another analogy. You can go to a library to borrow a book, but you can't go to the library and copy all the books for your own use.

I used to go to the library, find books with the relevant chapters related to what I wanted to learn, and the librarian would photo copy all the pages I wanted to take home. So I guess technically you could copy all the books for your own use. It's just impractical to photocopy every page of every book in a library.

When I was in libraries you could photocopy a percentage of a book (15% maybe?), although I doubt it was enforced. One could do many trips, but it is impractical, as you say.

Re: GPTBot – OpenAI’s Web Crawler

#125

I wonder how much the regression of ChatGPT is due to it adding new content which has its origin from ChatGPT. The blog and SEO spam with ChatGPT fluff is going through the roof, eventually all of that will get crawled too and the model will just get positively reinforced on its own output. Or is that not a concern?

0.1% chance

My reasons are:

- I don't recall seeing any evidence that OpenAI has included new data in pretraining beyond the previous limit (Sept. 2021?) for GPT-3.5 or GPT-4

- Maybe they did finetuning or RLHF on new data but this is likely to be highly curated data

- AI generated content should be absolutely tiny in comparison to the data they are already working with.

Re: GPTBot – OpenAI’s Web Crawler

#126
post #57
post #16

Earlier quoted context omitted.

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…

This is a very thought provoking point and it throughly stimulated me to think deeper and through. Purpose of my website is threefold, document my own knowledge, maybe some vanity and the urge to give back something to "someone" make a better living or similar. Things get interesting at corporate scale. There are fat VC funds, executives, board of directors and what not - making far more and far more comfortable than…

Yes, it is interesting. To me, the important thing is that our labour is exploited in many more (and many more malicious) ways than making an LLM 0.000001% better, maybe (or maybe it makes it worse!). Therefore, the problem isn't the AI, it is this giant financial machine which sucks value out of all who actually produce it, no matter what tools it uses to do so.

Re: GPTBot – OpenAI’s Web Crawler

#127
post #89

Earlier quoted context omitted.

A human reads a website, watches/clicks on ad, buys merch, subscribes to website, sets a bookmark to a website, shares an article, invites others etc. What will be the point of sharing knowledge or content if it will no longer be associated to an individual or organization?

Some would argue that it allows for cutting edge research that could potentially massively benefit humanity. Whether or not that pans out is to be determined, but that is the stated goal. The point then, as advocates for this technology believe is for the betterment of civilization, that by training a neural network on this knowledge will make that knowledge more accessible to the rest of us. One could of course deba…

Nobody is doubting that accessible and correct information is good for humanity, I am questioning how will that affect knowledge/content providers.

I repeat, what will be the motivation of an individual to share or provide valuable information if you decrease or eliminate any control of where and how that information appears?

Re: GPTBot – OpenAI’s Web Crawler

#128
post #31

Earlier quoted context omitted.

My understanding is that it's not easy to untrain a model of data already fed to it. Regarding noai tags - is this respected or just wishful?

Like every time you put content on the internet: you depend on their good will to respect these tags, or robots.txt. OpenAI can decide to ignore it. It's wishful thinking.

The next version of GPT might have better citations, and they could just refuse to cite things they were not allowed to crawl.

However, it's trivial to know whether the bot crawled your site or stopped at robots.txt.

Re: GPTBot – OpenAI’s Web Crawler

#129
post #15

Earlier quoted context omitted.

Depends on what the courts say. We'll have to see.

OpenAI would love that kinda regulation, it would basically kill free models.

Gotta pull the ladder up after you if you really want to maximise profits.

Re: GPTBot – OpenAI’s Web Crawler

#130
post #30

Earlier quoted context omitted.

The legal cases don't mean anything. The rule of law has all but disappeared from the corporate world. The idea that courts or regulators will be able to control AI is laughable. They are too corrupt, and they are way too slow.

I think a key idea is that with the amount of jurisdictions and number of courts the odds that a clean and sympathetic judge can be found approach one. I would argue that European jurisdictions are inherently less likely to be in pockets of American corporate interest and they are more likely to hear cases where fundamental human freedoms are at stake because both of these are existential threats to European independ…

The corporations that provide AI hold all the power because people (and businesses!) want to use their products.

Let's say the French government decides that OpenAI must change something about their business practices if they want to continue operating in France. OpenAI says "nope", and blocks access to French users.

Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countries can. How long do you think it will take before a storm of corporate outrage forces the government to relent?

Any individual government (except, perhaps, the combined US and EU governments) is powerless against today's technology megacorporations, because they can take much more away from a country than that country can take from them. If push ever comes to shove, it will become obvious where the true power lies. So far, the corporations have barely even tried to throw their weight around.

Post reply on HN