Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

191–200 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#191
post #119
post #4

Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT(?)

It’s so now they can lobby for anti scraping regulation and hamper any possible catch-up.

How's that gonna work when they need to update their model? Also, how would they compete with companies like FB that have an insane amount of conversational data, or Google, a company that literally indexes the internet?

Re: GPTBot – OpenAI’s Web Crawler

#192
post #186

Earlier quoted context omitted.

Which benefits the company.

Yeah but I don't think it's an inherently bad thing. +100M people use ChatGPT without paying anything, in this it benefits much more than the company.

Oh but they do pay. They pay their own time to gradually train the model and feed their data. There's no such thing as "free".

Re: GPTBot – OpenAI’s Web Crawler

#193
post #119

Earlier quoted context omitted.

It’s so now they can lobby for anti scraping regulation and hamper any possible catch-up.

How's that gonna work when they need to update their model? Also, how would they compete with companies like FB that have an insane amount of conversational data, or Google, a company that literally indexes the internet?

Spend money on licensing deals, lock out the competition. The value of the LLM isn’t up to date data, it’s the concepts of extracts. There’s very limited value in a large amount of crap if chinchilla is to be believed.

I don’t think stack overflow is all that valuable once your model has access to github due to their good friends at MS.

The money in proprietary AI is on the top end now, open source / edge is destroying monetisation on the lower end. Top end means high quality domain specific data.

Re: GPTBot – OpenAI’s Web Crawler

#194
post #186

Earlier quoted context omitted.

Yeah but I don't think it's an inherently bad thing. +100M people use ChatGPT without paying anything, in this it benefits much more than the company.

Oh but they do pay. They pay their own time to gradually train the model and feed their data. There's no such thing as "free".

That's just a win-win situation, you're using their services for free because it helps you, they use your interaction to improve the model; the model is still free to use.

Re: GPTBot – OpenAI’s Web Crawler

#195
post #169
post #130

Earlier quoted context omitted.

The corporations that provide AI hold all the power because people (and businesses!) want to use their products. Let's say the French government decides that OpenAI must change something about their business practices if they want to continue operating in France. OpenAI says "nope", and blocks access to French users. Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countri…

> Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countries can. How long do you think it will take before a storm of corporate outrage forces the government to relent? Bof, les alternatives à ChatGPT ne sont pas si mal. And even if the open source alternatives were far behind rather than just a bit — all this talk about corporate moats and their absence may be blind to t…

> Bof, les alternatives à ChatGPT ne sont pas si mal.

But that's not true, and people know it.

> the storms of protest in France are normally by the people, not by the corporations

Correct. CEOs of big corporations just call the ministers directly and tell them to get in line, or else.

Re: GPTBot – OpenAI’s Web Crawler

#196
post #189

Earlier quoted context omitted.

That would be a hilariously bad idea for them. Their business is based on fair use. The only way to enforce restrictions against scraping is through copyright law because obviously you can run the spidering code from any jurisdiction you want, so any law that says “thou shall not scrape” is toothless unless it acts through copyright. Any workable restrictions against using scraped data would also make ChatGPT illegal…

Nonsense. Regulation rarely works retroactively. Their model is trained and they have the money to license incremental data going forward, potentially exclusively.

Copyright laws do in fact (or have in fact) acted retroactively.

Re: GPTBot – OpenAI’s Web Crawler

#197
post #80

Earlier quoted context omitted.

They are business, not some random visitor. Google scrapes websites all the time, provide Adsense, a way to earn money.

Adsense pays you money for showing ads to human visitors, you don't get paid for allowing their crawler.

Google's crawler makes your page show up in their search results and gives you visitors.

Re: GPTBot – OpenAI’s Web Crawler

#198
post #30

Earlier quoted context omitted.

The legal cases don't mean anything. The rule of law has all but disappeared from the corporate world. The idea that courts or regulators will be able to control AI is laughable. They are too corrupt, and they are way too slow.

The legal cases don't "mean anything" because AI training is /legal/, not because courts are "corrupt". If anything is transformative, an AI that doesn't memorize its input is.

It doesnt memorize anything. It just needs gazillion parameters that approach the size of the training set to finesse its conversational accent.

Re: GPTBot – OpenAI’s Web Crawler

#199
post #113
post #40

Earlier quoted context omitted.

It's not about who learned what from whom, it's about superstar economy. If you serve all customers and leave nothing for the rest, it will be a problem. What do you think why writers and actors have included AI in the reasons of their strike?

> What do you think why writers and actors have included AI in the reasons of their strike? Because they are about to become obsolete, and they believe that screaming as loudly as they can is going to stop that. Their chances of success are roughly the same as if they were protesting against the law of gravity.

The fun part about being a strong believer of AI and actually understanding its capacities is being able to tell when people are completely blinded by hype.

AI will not make writers “obsolete”, that is utterly absurd. Would you say reality TV made tv writers obsolete? No? Oh well.

You get what you pay for. That includes what you pay for as a producer…

Re: GPTBot – OpenAI’s Web Crawler

#200

I wonder what kind of mischief facts people are going to start sneaking into OpenAI's newer models, by selectively feeding different responses to OpenAI when their crawler is identified.

Neat idea! The PR industrial complex has been trying so hard to convince us that the all-knowing all-seeing almighty AI is going to take our jobs and turn us into Soylent or whatever. Now let’s feed it some garbage and see if in all its glory it can tell sense from nonsense.

Plot twist: it can, like in the joke where the doctor tells the man who mixed all his family’s urine and his dog’s what they have
Post reply on HN