Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

71–80 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#71
post #62
post #19

Just a random thought: While they are already earning money by scraping data, would it not be nice if they pay the site owners a certain amount of the money they earn?

God no, can you imagine what the web will look like if every grifter can get paid just for existing?

Then you will find people who will share low-value (but easy to create content) that is used everywhere, a bit like these "isOdd, "isEven" NPM packages.

Re: GPTBot – OpenAI’s Web Crawler

#72
post #69

Friendship ended with SEO. Now LEO [1] is my best friend. [1]: LLM Engine Optimization

I don't think the word Engine belongs in there. It's just LLMO, LMAO.

Well the Engine could mean the Crawler part. Everyone is overloading technical terms for VC cash so :D

Re: GPTBot – OpenAI’s Web Crawler

#73
post #29
post #4

Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT(?)

GPT-4 finished training in August 2022, before the release of ChatGPT. If they had announced this sooner hardly anyone on the internet would have noticed. Props to them for adding it now.

Maybe some people weren't aware, but GPT-3 (and GPT-2, before that) APIs had been around for some time when ChatGPT was launched. I joined the private beta in early 2021.

Re: GPTBot – OpenAI’s Web Crawler

#74
post #67

Is there any argument in favor of commercial websites allowing GPTBot to crawl them? It's not like Google where allowing crawling brings you traffic. In fact, it's pretty much the opposite.

I assume by "commercial websites" you mean specifically "websites whose whole purpose is to have information in the website that you view in return for running ads?" Generally speaking, if I have a website where I have information about my business, then most likely it benefits me for people's LLMs to know that information, for the same reason I might buy an ad for my business.

Re: GPTBot – OpenAI’s Web Crawler

#75

What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…

If you're a marketer you won't care about citations. Just spam your product enough so that GPT "learns" it's the correct choice.

Re: GPTBot – OpenAI’s Web Crawler

#76
post #62
post #19

Just a random thought: While they are already earning money by scraping data, would it not be nice if they pay the site owners a certain amount of the money they earn?

God no, can you imagine what the web will look like if every grifter can get paid just for existing?

are OpenAI grifters? When Google can pay, why not them?

Re: GPTBot – OpenAI’s Web Crawler

#77
post #9

Earlier quoted context omitted.

quick back-of-envelope/googling: openai is worth $29,000,000.00 you contributed 0.00000000001 punches numbers in calculator thus the value of your free credits is 0.001 cents. minus any accounting fees.

>openai is worth $29,000,000.00 You might have missed a few zeros.

He doesn't believe they have a moat, seemingly!

Re: GPTBot – OpenAI’s Web Crawler

#78
post #30

Earlier quoted context omitted.

Hoping this is what they’ll use to train future models and deprecate the older ones before the legal cases proceed any further.

The legal cases don't mean anything. The rule of law has all but disappeared from the corporate world. The idea that courts or regulators will be able to control AI is laughable. They are too corrupt, and they are way too slow.

I think a key idea is that with the amount of jurisdictions and number of courts the odds that a clean and sympathetic judge can be found approach one. I would argue that European jurisdictions are inherently less likely to be in pockets of American corporate interest and they are more likely to hear cases where fundamental human freedoms are at stake because both of these are existential threats to European independence. In the US similar arguments can be made in states vs federal or the various federal circuits.

Courts are more deliberate than you would like — no denying that. But this is a feature not a flaw. It may be that damage will be done by then. Perhaps irreversible. But I would like to think if there is a will there is a way and that if things are terrible enough the governments will be bold in their responses.

Re: GPTBot – OpenAI’s Web Crawler

#79
post #66

Friendship ended with SEO. Now LEO [1] is my best friend. [1]: LLM Engine Optimization

What's the end goal? To teach it some very specific information, like about your company?

> To teach it some very specific information, like about your religion, nation-state, political party, controversial historic event,...

Re: GPTBot – OpenAI’s Web Crawler

#80
post #25
post #19

Just a random thought: While they are already earning money by scraping data, would it not be nice if they pay the site owners a certain amount of the money they earn?

If a human read a website and profited from the knowledge obtained, I don’t think we’d expect them to pay royalties to the site owner.

They are business, not some random visitor. Google scrapes websites all the time, provide Adsense, a way to earn money.
Post reply on HN