Live data from Hacker News

GPTBot – OpenAI’s Web Crawler

platform.openai.com

231–240 of 327 posts

Re: GPTBot – OpenAI’s Web Crawler

#231

Earlier quoted context omitted.

There's no win win situation. My content is stolen and given to others. I've lost. Google paid me for traffic via ads, therefore I allowed google to ingest my content. You as a person could read it. I've never given you permission to resell it, and if you did, I'd come after you to pay royalties. The same must apply to openai and other leeches.

> My content is stolen Physical property is stolen. Information is copied.

The term depends on use.

Physical property is either borrowed, owned, sold, and so on.

If your spouse takes your car to work without your knowledge it's borrowed. If they take it and sell it without consent it's theft.

Same applies to data. But data is electrons and as such it can't be moved, it is "copied". So technically speaking you are right, but practically you are not. If you steal NBC's prerelease movie then that's theft. As is copying it without constent. Once you pay for it you can copy it from their servers to your device. But you can't copy it to someone else's machine.

Re: GPTBot – OpenAI’s Web Crawler

#232

Earlier quoted context omitted.

Copyright laws do in fact (or have in fact) acted retroactively.

When? Not doubting, just curious about scope and type of scenarios where it's happened.

I cannot for the life of me find the links but I feel like this happened with Monopoly or some other board game.

Re: GPTBot – OpenAI’s Web Crawler

#233
post #193

Earlier quoted context omitted.

Spend money on licensing deals, lock out the competition. The value of the LLM isn’t up to date data, it’s the concepts of extracts. There’s very limited value in a large amount of crap if chinchilla is to be believed. I don’t think stack overflow is all that valuable once your model has access to github due to their good friends at MS. The money in proprietary AI is on the top end now, open source / edge is destroyi…

> The value of the LLM isn’t up to date data As a heavy ChatGPT user I disagree. Lack of up to date data is one of the biggest issues I face every day - technology changes fast, libraries change APIs, new tech comes out, etc.

I’m working on this problem (heavy user of chatgpt too). What kinds of libraries do you use it for that are out of date. I could hopefully get you into the beta with it having better responses for those libs. Please email me gaurav@gvkhna.com

Re: GPTBot – OpenAI’s Web Crawler

#234

I wonder what kind of mischief facts people are going to start sneaking into OpenAI's newer models, by selectively feeding different responses to OpenAI when their crawler is identified.

They could in theory combat it by comparing results with a second crawler that uses a different User Agent.

If they were going to the amount of energy to do a 2nd crawl using a different user agent, then why bother advertising the user agent at all and just feed it the Chrome one like every other home-grown spider does

Re: GPTBot – OpenAI’s Web Crawler

#236
post #203

Earlier quoted context omitted.

The fun part about being a strong believer of AI and actually understanding its capacities is being able to tell when people are completely blinded by hype. AI will not make writers “obsolete”, that is utterly absurd. Would you say reality TV made tv writers obsolete? No? Oh well. You get what you pay for. That includes what you pay for as a producer…

> AI will not make writers “obsolete”, that is utterly absurd. Of course. And those so-called "computers" won't make human calculators obsolete. After all, they are as large as an entire room, and by the time they are ready to receive input, a human with his slide rule has already computed three and a half entire logarithms! Human creative professions have 5-10 years left, if they are very lucky.

> Human creative professions have 5-10 years left, if they are very lucky.

So in that sense do developers have ~2 years left? Code is much more rigid than acting or creative writing and AI seems to be getting there first. I mean if the all-powerful AI can make modern movies than clearly it can handle writing all code right?

Re: GPTBot – OpenAI’s Web Crawler

#237
post #233

Earlier quoted context omitted.

> The value of the LLM isn’t up to date data As a heavy ChatGPT user I disagree. Lack of up to date data is one of the biggest issues I face every day - technology changes fast, libraries change APIs, new tech comes out, etc.

I’m working on this problem (heavy user of chatgpt too). What kinds of libraries do you use it for that are out of date. I could hopefully get you into the beta with it having better responses for those libs. Please email me gaurav@gvkhna.com

Rust libraries as well as Hashicorp Nomad (has changed a lot since ChatGPT's last training point). Also QuickWit is totally unknown to ChatGPT.

Re: GPTBot – OpenAI’s Web Crawler

#238
post #17

Earlier quoted context omitted.

This doesn't appear consistent with other visitors to your website. If a cafe owner uses info on your site to improve their baking, should they also be required to share their revenue with you?

If the cafe is a multi-billion corporation that can only exist because it can leech of content created by millions of other people without providing anything at all to them in return (and I'm not necessarily talking about financial compensation) then yeah.. maybe you should.

So Starbucks? Should they be sharing all their revenue with whoever invented all those Italian coffee drinks?

Re: GPTBot – OpenAI’s Web Crawler

#239

Earlier quoted context omitted.

> My content is stolen Physical property is stolen. Information is copied.

The term depends on use. Physical property is either borrowed, owned, sold, and so on. If your spouse takes your car to work without your knowledge it's borrowed. If they take it and sell it without consent it's theft. Same applies to data. But data is electrons and as such it can't be moved, it is "copied". So technically speaking you are right, but practically you are not. If you steal NBC's prerelease movie then t…

> If you steal NBC's prerelease movie then that's theft.

No. Advocates of expanded IP law have attempted to spread the idea that copyright infringement is "theft" as it adds emotional weight to their arguments. "You wouldn't download a car" etc. Same for the use of the word "piracy" - borrow an emotionally laden term from another context and hope nobody notices the sleight of hand.

And it's important that we reject this definition because it distorts the reality of the situation.

Re: GPTBot – OpenAI’s Web Crawler

#240
post #209

Earlier quoted context omitted.

That's of no benefit to me. Quite the contrary.

Why? If it helps people it should be good. Why bother posting something on the public web if not to help people. Sure a large org is receiving some ancillary benefit, but do you feel the same hostility for people working at [large corp] using what you worked on to help them at work? I honestly don't understand the hostility towards llms using public data

One of the reasons is that the company can later close up the effort, completely destroying the future potential helping part of it.

But at the end of the day, I understand that altruism doesn't work this way. But this just means that while I have some tendencies, I'm not altruistic after all. I attach a lot of feelings to where my work ends up and how it affects things, which is, for example, why I like "sticky" licenses like the GPL, and tend toward efforts like the Effective Altruism, however ineffective I think they end up being.

>I honestly don't understand the hostility towards llms using public data

So, getting back to the topic, feelings are attached to where the publications end up and how it affects things. Because of the unintended consequence of companies training AI on publicly available data, people harboring these feelings feel like their thing has been taken from them without their consent. And that is a bad feeling, powerless, inability, and one of the ways of coping with that is coping with it on the outside, directing the feeling outward, whereby it becomes active defense, or hostility.

Post reply on HN