Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

141–150 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#141
Man what a time we live in :) It's like history is being written (ok compiled & backpropagated) right under our feet !

I can see a bots.txt entry in the near future that discern the site's data-usage for bots vs humans

  User-agent-class: AI
  Data-Policy-Allow:  /news/* /articles/*
  Data-Policy-Deny:  */comments

Re: GPTBot – OpenAI’s Web Crawler

#142
post #113
post #40

Earlier quoted context omitted.

It's not about who learned what from whom, it's about superstar economy. If you serve all customers and leave nothing for the rest, it will be a problem. What do you think why writers and actors have included AI in the reasons of their strike?

> What do you think why writers and actors have included AI in the reasons of their strike? Because they are about to become obsolete, and they believe that screaming as loudly as they can is going to stop that. Their chances of success are roughly the same as if they were protesting against the law of gravity.

Not as obsolete as their bosses/owners have long been.

We make our own rules. We decide what to allow and what to value. If technology changes something, it's because we let it.

Re: GPTBot – OpenAI’s Web Crawler

#143
post #30

Earlier quoted context omitted.

Hoping this is what they’ll use to train future models and deprecate the older ones before the legal cases proceed any further.

The legal cases don't mean anything. The rule of law has all but disappeared from the corporate world. The idea that courts or regulators will be able to control AI is laughable. They are too corrupt, and they are way too slow.

The legal cases don't "mean anything" because AI training is /legal/, not because courts are "corrupt". If anything is transformative, an AI that doesn't memorize its input is.

Re: GPTBot – OpenAI’s Web Crawler

#144
post #102

Earlier quoted context omitted.

If you think copyright lawyers and the entertainment industry is going to let some AI upstarts launder their IP without a fight you aren't paying attention.

> AI upstarts You mean corporations that wield more power than most governments, and have revenues equivalent to the GDP of entire countries? If Universal or 20th Century Fox were to ever become a serious obstacle, Google and Microsoft are simply going to buy them. This isn't the early 2000s anymore. The power balance has shifted dramatically .

FAANG already haven't bought or started competitors to the record labels they resell in their music stores. Don't see why they'll start now.

Re: GPTBot – OpenAI’s Web Crawler

#145
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

What if your website contains incorrect information that makes their model worse?

Ahh DarkHatAI-Patterns almost like DarkHat SEO-Techniques.

Re: GPTBot – OpenAI’s Web Crawler

#146
post #16

Earlier quoted context omitted.

Counterpoint (not just to be annoying — I think you pose a very interesting unanswered question): If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me? I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff…

Interesting point though I'd go with another analogy. You can go to a library to borrow a book, but you can't go to the library and copy all the books for your own use.

You can do that. Google literally already did that.

Re: GPTBot – OpenAI’s Web Crawler

#147
post #38
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

scooba diving This is a bit pedantic but the term is "scuba diving". Scuba is an acronym that's short for "self contained underwater breathing apparatus". It doesn't work if you don't spell it right.

ChatGPT bot detected

Re: GPTBot – OpenAI’s Web Crawler

#148
post #30

Earlier quoted context omitted.

The legal cases don't mean anything. The rule of law has all but disappeared from the corporate world. The idea that courts or regulators will be able to control AI is laughable. They are too corrupt, and they are way too slow.

The legal cases don't "mean anything" because AI training is /legal/, not because courts are "corrupt". If anything is transformative, an AI that doesn't memorize its input is.

Yet gleefully emits its training data when one asks the right questions. It can be code, prose or images.

Yeah, doesn't remember. Mhm...

Oh, it just can't remember the license terms of the code it "reads", so it can't comply with these licenses or help people to comply with these licenses.

Convenient.

Re: GPTBot – OpenAI’s Web Crawler

#149
post #45

Earlier quoted context omitted.

Every response so far is no i.e. the hobby website doesn't merit any compensation. A contrarian take to support the original commenter is that if the site owner had ads, i probably got him or her some increment in site visits and helped in some small way with monetization, site ranking and boosted his / her public persona, credibility. When GPT bot visits, none of that happens. Much worse - people who might have visi…

The LinkedIn case has already established the legality of scraping, so this argument falls flat too.

[deleted]

Re: GPTBot – OpenAI’s Web Crawler

#150
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

[deleted]
Post reply on HN