Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

241–250 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#241
post #4

Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT(?)

At least now you can see if your website is being crawled by them. It also exposes them to be easily targeted to send them invalid data or even misinformation. Before people will already doing that before by putting information that people wouldn’t see like white text on a white background.

Re: GPTBot – OpenAI’s Web Crawler

#242

I wonder what kind of mischief facts people are going to start sneaking into OpenAI's newer models, by selectively feeding different responses to OpenAI when their crawler is identified.

I did this to Google for a while only to have my domains listed as malicious. I did not offer any malicious material, just different content for search engines was enough to flag my sites. They also did this to me when I gave google different IP addresses using a split DNS view. This was a while back so maybe they stopped this, I honestly don't know. Now I just give them and most bots a password prompt. Google and mo…

Dumb question but how would they know the content is different unless they're also crawling incognito and comparing the results?

Re: GPTBot – OpenAI’s Web Crawler

#243
post #194

Earlier quoted context omitted.

Oh but they do pay. They pay their own time to gradually train the model and feed their data. There's no such thing as "free".

That's just a win-win situation, you're using their services for free because it helps you, they use your interaction to improve the model; the model is still free to use.

It would be a win-win if the company promised that they'll keep the AI as it is, and as free as it is, as long as the company functions. Then they would take something, give something, and we could discuss if what we get outweighs what they took.

But the street is one-way, and it's the company that has the upper hand. The company can (and does) retract access to the AI, but they themselves keep what they took. If in the meantime people became attached to what the company gave, the company even does damage to them, not just by taking away the access, but because of severing the supply for a dependency.

So the people are taken advantage of because the company took the assets, they are taken advantage of because they help to further train the AI by using it, and then they get, at most, the privilege to pay for something that grew out of them.

That's why it's not a win-win. It's a win for the company, and a questionable outcome, and a risk for the people.

Re: GPTBot – OpenAI’s Web Crawler

#244
post #17
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

This doesn't appear consistent with other visitors to your website. If a cafe owner uses info on your site to improve their baking, should they also be required to share their revenue with you?

I see no point in engaging with this argument. ChatGPT is not a human. I, nor anyone else, should have to explain to you why that makes all the difference here.

Re: GPTBot – OpenAI’s Web Crawler

#245

Earlier quoted context omitted.

I did this to Google for a while only to have my domains listed as malicious. I did not offer any malicious material, just different content for search engines was enough to flag my sites. They also did this to me when I gave google different IP addresses using a split DNS view. This was a while back so maybe they stopped this, I honestly don't know. Now I just give them and most bots a password prompt. Google and mo…

Dumb question but how would they know the content is different unless they're also crawling incognito and comparing the results?

Keeping you honest with incognito crawling is something they have to do anyway, to catch various tricks and scams - malware served up to users, etc.

Re: GPTBot – OpenAI’s Web Crawler

#246
post #89

Earlier quoted context omitted.

Some would argue that it allows for cutting edge research that could potentially massively benefit humanity. Whether or not that pans out is to be determined, but that is the stated goal. The point then, as advocates for this technology believe is for the betterment of civilization, that by training a neural network on this knowledge will make that knowledge more accessible to the rest of us. One could of course deba…

Nobody is doubting that accessible and correct information is good for humanity, I am questioning how will that affect knowledge/content providers. I repeat, what will be the motivation of an individual to share or provide valuable information if you decrease or eliminate any control of where and how that information appears?

> I repeat, what will be the motivation of an individual to share or provide valuable information if you decrease or eliminate any control of where and how that information appears?

Forum users don’t seem to mind. Reddit, HN, Twitter, Facebook, etc. are all examples of users freely providing valuable content without expectation or control.

I suppose it’s also not too different from a listener summarizing a speech. When you speak publicly you don’t get to control who hears it or how they will interpret it.

Re: GPTBot – OpenAI’s Web Crawler

#247

Earlier quoted context omitted.

> What’s the incentive for people to allow the crawler at all? So that LLMs can learn from it? Profit is not the only thing that motivates people. I’ve spent years contributing to Stack Overflow to help people solve their problems, with the understanding that they had an open data policy and anybody could access the data dump easily to build things with it. It pisses me off that they are now trying to lock that infor…

By the same token (no pun intended), locking up such data in a closed (in many senses) LLM wouldn’t be a desirable outcome?

How does an LLM learning from an open dataset lock it up?

Re: GPTBot – OpenAI’s Web Crawler

#248
It's painfully clear that the future of the web is not open. Everyone is going to put their site behind a paywall with onerous TOS to keep AI companies from stealing all their content. RIP the golden age of the internet.

Re: GPTBot – OpenAI’s Web Crawler

#249

Earlier quoted context omitted.

The term depends on use. Physical property is either borrowed, owned, sold, and so on. If your spouse takes your car to work without your knowledge it's borrowed. If they take it and sell it without consent it's theft. Same applies to data. But data is electrons and as such it can't be moved, it is "copied". So technically speaking you are right, but practically you are not. If you steal NBC's prerelease movie then t…

> If you steal NBC's prerelease movie then that's theft. No. Advocates of expanded IP law have attempted to spread the idea that copyright infringement is "theft" as it adds emotional weight to their arguments. "You wouldn't download a car" etc. Same for the use of the word "piracy" - borrow an emotionally laden term from another context and hope nobody notices the sleight of hand. And it's important that we reject t…

> And it's important that we reject this definition because it distorts the reality of the situation.

Depends who's reality. A content creator's reality is that their content is indeed stolen and monetised by someone without permission.

"Advocates of expanded IP law" do appear to be in the right, at least by law. Copying and distributing digital products is treated more or less as theft, particularly when done at scale.

AI and current training practices are even worse than stealing someone's work. It steals someone's identity. AI can copy unique characteristics, not just individual content to reproduce identical content. It can replicate a person's unique style without consent, and that's uniquely dangerous.

Re: GPTBot – OpenAI’s Web Crawler

#250
post #180

Earlier quoted context omitted.

>so there’s no benefit in allowing them access. Well, you're helping improving the model.

That's of no benefit to me. Quite the contrary.

If I read and learn from your content it's of no benefit to you either.

If you don't want others to learn from what you have to say, just talk to a brick wall.

Post reply on HN