Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

251–260 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#251
post #21
post #5

If you scrape my hobby website about photography, scuba diving or let's say baking or gardening which improves your model by let's say a delta of 0.00000000001 than shouldn't I get some free credits to use that model or proportionate share in the revenue stream? EDIT: scuba diving NOT scooba diving

If I learn something from your StackOverflow answers, do you expect me to share a percentage of my future salary with you?

SO answers are explicitly licensed under CC-BY-SA 2.5/3/4, depending on the time it was posted https://stackoverflow.com/help/licensing. So no.

Re: GPTBot – OpenAI’s Web Crawler

#252
post #194

Earlier quoted context omitted.

That's just a win-win situation, you're using their services for free because it helps you, they use your interaction to improve the model; the model is still free to use.

It would be a win-win if the company promised that they'll keep the AI as it is, and as free as it is, as long as the company functions. Then they would take something, give something, and we could discuss if what we get outweighs what they took. But the street is one-way, and it's the company that has the upper hand. The company can (and does) retract access to the AI, but they themselves keep what they took. If in…

It’s win-win based on current usage. Even if OpenAI got shut down, I still benefited from using it.

Many good things don’t last forever. If they go away that doesn’t invalidate the experiences you had.

Re: GPTBot – OpenAI’s Web Crawler

#253

Earlier quoted context omitted.

> If you steal NBC's prerelease movie then that's theft. No. Advocates of expanded IP law have attempted to spread the idea that copyright infringement is "theft" as it adds emotional weight to their arguments. "You wouldn't download a car" etc. Same for the use of the word "piracy" - borrow an emotionally laden term from another context and hope nobody notices the sleight of hand. And it's important that we reject t…

> And it's important that we reject this definition because it distorts the reality of the situation. Depends who's reality. A content creator's reality is that their content is indeed stolen and monetised by someone without permission. "Advocates of expanded IP law" do appear to be in the right, at least by law. Copying and distributing digital products is treated more or less as theft, particularly when done at sca…

> Depends who's reality

On a trivial level this is correct as words mean what we collectively decide they mean.

However I am making the point that a) the meaning has been changed and b) it has changed in a way that is deceptive and masks a useful fact about the world

Re: GPTBot – OpenAI’s Web Crawler

#254
post #4

Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT(?)

Their papers say they were using Common Crawl for crawling. If you didn't want your pages in Common Crawl (eg. Twitter didn't) for use in many downstream analyses or uses beyond just OA, you could already have said so in your robots.txt.

Re: GPTBot – OpenAI’s Web Crawler

#255
post #119
post #4

Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT(?)

It’s so now they can lobby for anti scraping regulation and hamper any possible catch-up.

They still need current data or their GPT models will be stuck at september 2021 forever

Re: GPTBot – OpenAI’s Web Crawler

#256

Earlier quoted context omitted.

The legal cases don't "mean anything" because AI training is /legal/, not because courts are "corrupt". If anything is transformative, an AI that doesn't memorize its input is.

It doesnt memorize anything. It just needs gazillion parameters that approach the size of the training set to finesse its conversational accent.

LLama2 has a 5TB training set.

Re: GPTBot – OpenAI’s Web Crawler

#257

Earlier quoted context omitted.

Dumb question but how would they know the content is different unless they're also crawling incognito and comparing the results?

Keeping you honest with incognito crawling is something they have to do anyway, to catch various tricks and scams - malware served up to users, etc.

So robots.txt is meaningless if they have to violate it to check for malicious content in blocked off pages anyways.

Re: GPTBot – OpenAI’s Web Crawler

#258

Earlier quoted context omitted.

Copyright laws do in fact (or have in fact) acted retroactively.

When? Not doubting, just curious about scope and type of scenarios where it's happened.

I'm going largely by memory but when the U.S. expanded copyright at one point they actually took some stuff out of the public domain. You can look it up but the current formula is authors life plus 70 and a different formula for corporate works, and when they expanded it most recently there were actually some public domain works that become not public domain retroactively. (A quick google search reveals the 1976 Act added 19 years to the terms of existing copyrights, this might be what I'm thinking of-- in other words some works that had copyright expired then had them renewed and removed from the public domain.)

There's also copyright reversion, which is a related new provision that applied to older copyrighted works. Quoting from an article I just pulled up

"...the 1976 Act created a new right allowing authors and their heirs to terminate a prior grant of copyright, the Act also set forth specific steps concerning the timing and contents of the termination notice that must be served in order to effectuate termination. The termination of a grant may be effective “at any time during a period of five years beginning of the end of 56 years from the date the copyright was originally secured”..."

But this is a red herring because the fact a model has been trained in the past doesn't mean a copyright lawsuit is "retroactive". The infringement would presumably be occuring anew every day you make it available on your web site.

Re: GPTBot – OpenAI’s Web Crawler

#259

Earlier quoted context omitted.

By the same token (no pun intended), locking up such data in a closed (in many senses) LLM wouldn’t be a desirable outcome?

How does an LLM learning from an open dataset lock it up?

I meant the LLM weights are not publicly available in the case of ØpenAI, so whatever you contribute to it will be locked up, just like SO locked up their user-generated data.

Re: GPTBot – OpenAI’s Web Crawler

#260
post #254
post #4

Nice of them to respect crawling after they've already trained their model. Presumably these headers don't affect any pages they've already crawled to train GPT(?)

Their papers say they were using Common Crawl for crawling. If you didn't want your pages in Common Crawl (eg. Twitter didn't) for use in many downstream analyses or uses beyond just OA, you could already have said so in your robots.txt.

Opt out != opt in. This reminds me of the beginning of the hitchhikers guide to the galaxy where Dent’s house is being demolished but the notice had been on display in a locked basement below city hall or something. He could have objected, technically!
Post reply on HN