Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

131–140 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#131
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

What if your website contains incorrect information that makes their model worse?

Re: GPTBot – OpenAI’s Web Crawler

#132

I wonder what kind of mischief facts people are going to start sneaking into OpenAI's newer models, by selectively feeding different responses to OpenAI when their crawler is identified.

Neat idea!

The PR industrial complex has been trying so hard to convince us that the all-knowing all-seeing almighty AI is going to take our jobs and turn us into Soylent or whatever. Now let’s feed it some garbage and see if in all its glory it can tell sense from nonsense.

Re: GPTBot – OpenAI’s Web Crawler

#133
"As an AI language model, I don't have personal opinions or preferences. However, I can provide some information based on my training data up to September 2021."

I'm confused.. if it's being trained on data up to a certain date, than why would the web crawler matter?

Re: GPTBot – OpenAI’s Web Crawler

#134
post #60

Earlier quoted context omitted.

Interesting point though I'd go with another analogy. You can go to a library to borrow a book, but you can't go to the library and copy all the books for your own use.

Not really sure that this analogy applies, because I could definitely photocopy as many books from the library as I physically can. No one is going to stop me.

Well it's not so much about the physical act of doing it, it's the trying to convince the world it's for your own private use and not for commercial gain.

Otherwise, intellectual property laws can perhaps apply.

It'd be a hard push to claim it's fair use, a wholesale copying of other's works.

Re: GPTBot – OpenAI’s Web Crawler

#136

"As an AI language model, I don't have personal opinions or preferences. However, I can provide some information based on my training data up to September 2021." I'm confused.. if it's being trained on data up to a certain date, than why would the web crawler matter?

Paraphrase: However, I can provide some information up to September 2021 based on my training data

I believe the chatbot is prompted to not answer for things after Sept 2021, rather than the data itself being limited.

I could be wrong though.

Re: GPTBot – OpenAI’s Web Crawler

#137
post #80
post #25

Earlier quoted context omitted.

If a human read a website and profited from the knowledge obtained, I don’t think we’d expect them to pay royalties to the site owner.

They are business, not some random visitor. Google scrapes websites all the time, provide Adsense, a way to earn money.

Adsense pays you money for showing ads to human visitors, you don't get paid for allowing their crawler.

Re: GPTBot – OpenAI’s Web Crawler

#138

Earlier quoted context omitted.

A human reads a website, watches/clicks on ad, buys merch, subscribes to website, sets a bookmark to a website, shares an article, invites others etc. What will be the point of sharing knowledge or content if it will no longer be associated to an individual or organization?

Even something as simple as a blog accrues the author some small reputational benefit. Of course if people do conclude that there’s no benefit to sharing knowledge and stop doing so then those who do share knowledge will have an outsized impact on AI training. In the extreme: the opportunity to create truth. Thus the incentive to publish in order to stop those people is created.

Oh gee, opportunity for volunteers to create truth versus resourceful companies and organisation eager to share their own version of truth on a system whose workings barely anyone can comprehend and for-profit company whose actions anyone can predict.

Sorry for the sarcasm, but your comment is essentially "Lets remove motivation for those who had incentive to share valuable information and see how it turns out.".

Re: GPTBot – OpenAI’s Web Crawler

#139
post #130

Earlier quoted context omitted.

I think a key idea is that with the amount of jurisdictions and number of courts the odds that a clean and sympathetic judge can be found approach one. I would argue that European jurisdictions are inherently less likely to be in pockets of American corporate interest and they are more likely to hear cases where fundamental human freedoms are at stake because both of these are existential threats to European independ…

The corporations that provide AI hold all the power because people (and businesses!) want to use their products. Let's say the French government decides that OpenAI must change something about their business practices if they want to continue operating in France. OpenAI says "nope", and blocks access to French users. Suddenly French companies aren't able to use GPT-X anymore – while their competitors in other countri…

> Let's say the French government decides that OpenAI must change something about their business practices if they want to continue operating in France. OpenAI says "nope", and blocks access to French users.

That's one possible outcome. (ETA: You DO have a point here, but...)

The other is, you know, something like every website explicitly telling me, via an annoying popup, how much they value my privacy. Also, me not being able to access half of US news sites to this day.

The last time EU raised their finger, every technology company (FAANG included) shat their pants.

And that was simpler times, times when a cookie stored in your temp folder without websites shouting they're about to do so, was somehow the biggest concern of an EU netizen. It almost seems ridiculous, compared to the damage AI could do (the extent of which which nobody really knows).

Re: GPTBot – OpenAI’s Web Crawler

#140
post #17

Earlier quoted context omitted.

This doesn't appear consistent with other visitors to your website. If a cafe owner uses info on your site to improve their baking, should they also be required to share their revenue with you?

Are you considering ad revenue?

That already puts the website in the sleazy category. "I mixed my helpful information with mind poison" isn't a strong position to argue fair play from.
Post reply on HN