Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

141–150 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#141
post #86

Earlier quoted context omitted.

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…

With Wittgenstein I think we see that "hallucinations" are a part of language in general, albeit one I could see being particularly vexing if you're trying to build a perfectly controllable chatbot.

This sounds interesting, could you give more detail on what you're referring to?

Re: Anyone got a contact at OpenAI. They have a spider problem

#142
post #106

Earlier quoted context omitted.

And they do have (the same) robots.txt on every domain, tailored for GPTbot, i.e. https://petra-cody-carlene.web.sp.am/robots.txt So, GPTBot is not following robots.txt, apparently.

Accessing a directly referenced page is common in order to receive the noindex header and/or meta tag, whose semantics are not implied by “Disallow: /” And then all the links are to external domains, which aren't subject to the first site's robots.txt

This is a moderately persuasive argument.

Although the crawler should probably ignore all the html body. But it does feel like a grey area if I accept your first pint.

Re: Anyone got a contact at OpenAI. They have a spider problem

#143
post #65

Earlier quoted context omitted.

He says 3 million, and 1.8 million are for robots.txt So 1.2 million non robots.txt requests, when his robots.txt file is configured as follows # buzz off User-agent: GPTBot Disallow: / Theoretically if they were actually respecting robots.txt they wouldn't crawl any pages on the site. Which would also mean they wouldn't be following any links... aka not finding the N subdomains.

A lot of crawlers, if not all, have a policy like "if you disallow our robot, it might take a day or two before it notices". They surely follow the path "check if we have robots.txt that allows us to scan this site, if we don't get and store robots.txt, scan at least the root of the site and its links". There won't be a second scan, and they consider that they are respecting robots.txt. Kind of "better ask for forgiv…

That is indistinguishable from not respecting robots.txt. There is a robots.txt on the root the first time they ask for it, and they read the page and follow its links regardless.

Re: Anyone got a contact at OpenAI. They have a spider problem

#145
post #15

I’d let them do their thing, why not?! They want the internet? This is the real internet. It looks like he doesn’t really care that much that they’re retrieving millions of pages, so let them do their thing…

Some scrapers respect robots.txt. OpenAI doesn't. SP is just informing the world at large of this fact.

Re: Anyone got a contact at OpenAI. They have a spider problem

#146
post #116

If they don't respect robots.txt then block them using a firewall or other server config. All of these companies are parasites.

The entire purpose of this website is to identify bad actors who do not respect robots.txt, so that they can be publicly shamed.

Re: Anyone got a contact at OpenAI. They have a spider problem

#147

With all the news about scraping legality you'd think a multi billion dollar AI company would try to obfuscate their attempts.

If you're not walling off your content behind a login that contains terms that you agree to not scraping, then, scraping that site is 100% legal. Robots.txt isn't a legal document.

I frequently respect the wishes of other people without any legal obligation to do so, in business, personal, and anonymous interactions.

I do try to avoid people that use the law as a ceiling for the extension of their courtesy to others, as they are consistently quite terrible people.

Re: Anyone got a contact at OpenAI. They have a spider problem

#148
post #137
post #86

Earlier quoted context omitted.

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…

Contrast with Q&A on products on Amazon where people routinely answer that way. I have flagged responses saying "I don't know" but nothing ever comes of it.

I’d place in the same category the responses that I give to those chat popups so many sites have. They show a person saying to me “Can I help you with anything today?” so I always send back “No”.

Re: Anyone got a contact at OpenAI. They have a spider problem

#149
post #71

Eventually, OpenAI (and friends) are going to be training their models on almost exclusively AI generated content, which is more often than not slightly incorrect when it comes to Q&A, and the quality of AI responses trained on that content will quickly deteriorate. Right now, most internet content is written by humans. But in 5 years? Not so much. I think this is one of the big problems that the AI space needs to so…

The end state of training on web text has always been an ouroboros - primarily because of adtech incentives to produce low quality content at scale to capture micro pennies. The irony of the whole thing is brutal.

Content you’re allowed and capable of scraping on the Internet is such a small amount of data, not sure why people are acting otherwise.

Common crawl alone is only a few hundred TB, I have more content than that on a NAS sitting in my office that I built for a few grand (Granted I’m a bit of a data hoarder). The fears that we have “used all the data” are incredibly unfounded.

Re: Anyone got a contact at OpenAI. They have a spider problem

#150

I'm more interested in what that content farm is for. It looks pointless, but I suspect there's a bizarre economic incentive. There are affiliate links, but how much could that possibly bring in?

It's for shits-and-giggles and it's doing its job really well right now. Not everything needs to have an economic purpose, 100 trackers, ads and backed by a company.
Post reply on HN