Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT(?)
It’s so now they can lobby for anti scraping regulation and hamper any possible catch-up.
GPTBot – OpenAI’s Web Crawler
191–200 of 327 posts
Re: GPTBot – OpenAI’s Web Crawler
#192Earlier quoted context omitted.
Which benefits the company.
Yeah but I don't think it's an inherently bad thing. +100M people use ChatGPT without paying anything, in this it benefits much more than the company.
Re: GPTBot – OpenAI’s Web Crawler
#193Earlier quoted context omitted.
It’s so now they can lobby for anti scraping regulation and hamper any possible catch-up.
How's that gonna work when they need to update their model? Also, how would they compete with companies like FB that have an insane amount of conversational data, or Google, a company that literally indexes the internet?
I don’t think stack overflow is all that valuable once your model has access to github due to their good friends at MS.
The money in proprietary AI is on the top end now, open source / edge is destroying monetisation on the lower end. Top end means high quality domain specific data.
Re: GPTBot – OpenAI’s Web Crawler
#194Earlier quoted context omitted.
Yeah but I don't think it's an inherently bad thing. +100M people use ChatGPT without paying anything, in this it benefits much more than the company.
Oh but they do pay. They pay their own time to gradually train the model and feed their data. There's no such thing as "free".
Re: GPTBot – OpenAI’s Web Crawler
#195Earlier quoted context omitted.
The corporations that provide AI hold all the power because people (and businesses!) want to use their products. Let's say the French government decides that OpenAI must change something about their business practices if they want to continue operating in France. OpenAI says "nope", and blocks access to French users. Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countri…
> Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countries can. How long do you think it will take before a storm of corporate outrage forces the government to relent? Bof, les alternatives à ChatGPT ne sont pas si mal. And even if the open source alternatives were far behind rather than just a bit — all this talk about corporate moats and their absence may be blind to t…
But that's not true, and people know it.
> the storms of protest in France are normally by the people, not by the corporations
Correct. CEOs of big corporations just call the ministers directly and tell them to get in line, or else.
Re: GPTBot – OpenAI’s Web Crawler
#196Earlier quoted context omitted.
That would be a hilariously bad idea for them. Their business is based on fair use. The only way to enforce restrictions against scraping is through copyright law because obviously you can run the spidering code from any jurisdiction you want, so any law that says “thou shall not scrape” is toothless unless it acts through copyright. Any workable restrictions against using scraped data would also make ChatGPT illegal…
Nonsense. Regulation rarely works retroactively. Their model is trained and they have the money to license incremental data going forward, potentially exclusively.
Re: GPTBot – OpenAI’s Web Crawler
#197Earlier quoted context omitted.
They are business, not some random visitor. Google scrapes websites all the time, provide Adsense, a way to earn money.
Adsense pays you money for showing ads to human visitors, you don't get paid for allowing their crawler.
Re: GPTBot – OpenAI’s Web Crawler
#198Earlier quoted context omitted.
The legal cases don't mean anything. The rule of law has all but disappeared from the corporate world. The idea that courts or regulators will be able to control AI is laughable. They are too corrupt, and they are way too slow.
The legal cases don't "mean anything" because AI training is /legal/, not because courts are "corrupt". If anything is transformative, an AI that doesn't memorize its input is.
Re: GPTBot – OpenAI’s Web Crawler
#199Earlier quoted context omitted.
It's not about who learned what from whom, it's about superstar economy. If you serve all customers and leave nothing for the rest, it will be a problem. What do you think why writers and actors have included AI in the reasons of their strike?
> What do you think why writers and actors have included AI in the reasons of their strike? Because they are about to become obsolete, and they believe that screaming as loudly as they can is going to stop that. Their chances of success are roughly the same as if they were protesting against the law of gravity.
AI will not make writers “obsolete”, that is utterly absurd. Would you say reality TV made tv writers obsolete? No? Oh well.
You get what you pay for. That includes what you pay for as a producer…
Re: GPTBot – OpenAI’s Web Crawler
#200I wonder what kind of mischief facts people are going to start sneaking into OpenAI's newer models, by selectively feeding different responses to OpenAI when their crawler is identified.
Neat idea! The PR industrial complex has been trying so hard to convince us that the all-knowing all-seeing almighty AI is going to take our jobs and turn us into Soylent or whatever. Now let’s feed it some garbage and see if in all its glory it can tell sense from nonsense.