Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

271–280 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#271

Earlier quoted context omitted.

They could in theory combat it by comparing results with a second crawler that uses a different User Agent.

If they were going to the amount of energy to do a 2nd crawl using a different user agent, then why bother advertising the user agent at all and just feed it the Chrome one like every other home-grown spider does

[deleted]

Re: GPTBot – OpenAI’s Web Crawler

#272

What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…

[deleted]

Re: GPTBot – OpenAI’s Web Crawler

#273
"Allowing GPTBot to access your site can help AI models become more accurate and improve their general capabilities and safety."

Well, "more accurate" means roughly: "So that your content can be absorbed and used as output by our generator."

Google at least linked to your website, while ChatGPT hides your website and only uses your content.

Re: GPTBot – OpenAI’s Web Crawler

#275

"Allowing GPTBot to access your site can help AI models become more accurate and improve their general capabilities and safety." Well, "more accurate" means roughly: "So that your content can be absorbed and used as output by our generator." Google at least linked to your website, while ChatGPT hides your website and only uses your content.

Google started to put answers before links a while ago. At least for the simple searches.

And now they also put many links of SEO spam with ads firsts.

Re: GPTBot – OpenAI’s Web Crawler

#276

Earlier quoted context omitted.

It would be a win-win if the company promised that they'll keep the AI as it is, and as free as it is, as long as the company functions. Then they would take something, give something, and we could discuss if what we get outweighs what they took. But the street is one-way, and it's the company that has the upper hand. The company can (and does) retract access to the AI, but they themselves keep what they took. If in…

It’s win-win based on current usage. Even if OpenAI got shut down, I still benefited from using it. Many good things don’t last forever. If they go away that doesn’t invalidate the experiences you had.

I agree wrt/ experience, but I don't think it applies to this situation. Even if you had an experience that would end, their ownership of the data wouldn't, and that, among other things, make this very one-sided.

I do want to stress something from your conclusion though. That people do better if they anticipate change, and can adapt to it.

Re: GPTBot – OpenAI’s Web Crawler

#277

"As an AI language model, I don't have personal opinions or preferences. However, I can provide some information based on my training data up to September 2021." I'm confused.. if it's being trained on data up to a certain date, than why would the web crawler matter?

For future models.

Re: GPTBot – OpenAI’s Web Crawler

#279
post #209

Earlier quoted context omitted.

That's of no benefit to me. Quite the contrary.

Why? If it helps people it should be good. Why bother posting something on the public web if not to help people. Sure a large org is receiving some ancillary benefit, but do you feel the same hostility for people working at [large corp] using what you worked on to help them at work? I honestly don't understand the hostility towards llms using public data

This is like asking why someone doesn't want to do free work for Oracle's database offerings. I mean, why not try to make things better?

Well, because a lot of corporations couldn't care less about the public good and are happy to cause harm if it makes them more money. OpenAI doesn't care about your welfare or mine any more than a sleezy ad company or spyware product does.

If OpenAI were actually an open source company working to benefit the broader ecosystem I would agree with you, but that's about as far as possible from the current state.

Re: GPTBot – OpenAI’s Web Crawler

#280
post #60

Earlier quoted context omitted.

Interesting point though I'd go with another analogy. You can go to a library to borrow a book, but you can't go to the library and copy all the books for your own use.

Not really sure that this analogy applies, because I could definitely photocopy as many books from the library as I physically can. No one is going to stop me.

As far as I'm aware, photocopying an entire book does in fact violate copyright law and librarians will refuse to help you do it: https://guides.cuny.edu/cunyfairuse/librarians
Post reply on HN