Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

271–280 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#271

Earlier quoted context omitted.

And they do have (the same) robots.txt on every domain, tailored for GPTbot, i.e. https://petra-cody-carlene.web.sp.am/robots.txt So, GPTBot is not following robots.txt, apparently.

All the lines related to GPTBot are commented out. That robots.txt isn't trying to block it. Either it has been changed recently or most of this comment thread is mistaken.

It wasn't commented out a few hours ago when I checked it. I think that's a recent change.

Re: Anyone got a contact at OpenAI. They have a spider problem

#272

Honeypots like this seem like a super interesting way to poison LLM training.

Any data scraped would be instantly deduplicated after the fact by whatever semantic dedupe engine they've cooked up.

What has it got to do with deduplication? I'm talking about crafting some kind of alternative (not necessarily duplicate) data. I agree some kind of post data collection cleaning/filtering of the data before training could potentially catch it. But maybe not!

Re: Anyone got a contact at OpenAI. They have a spider problem

#274
post #79

Earlier quoted context omitted.

Yeah Microsoft, I mean openai, compressing it into an ”ai voice” is “open”.

Hackernews really instant able to understand dark jokes are they.

You didn't really do the joke right. The second person is supposed to be doing the thing requested, but in a way the first person doesn't like. In this case, that would be AI companies contributing to a "free and open internet", but doing it """wrong""". But they're not contributing at all. The problem isn't that "free and open internet" is getting monkeys-pawed, it's all this closed proprietary stuff.

Re: Anyone got a contact at OpenAI. They have a spider problem

#275
post #200

Earlier quoted context omitted.

Here you go (1 req/min, 10 bytes/sec), please report results :) http { limit_req_zone $binary_remote_addr zone=ten_bytes_per_second:10m rate=1r/m; server { location / { if ($http_user_agent = "mimo") { limit_req zone=ten_bytes_per_second burst=5; limit_rate 10; } } } }

Scrapers of the future won't be ifElse logic, they will be LLM agents themselves. The slow loris robots.txt has to provide an interface to it's own LLM, which engages the scraper LLM in conversation, aiming to extend it as long as possible. "OK I will tell you whether or not I can be scraped. BUT FIRST, listen to this offer. I can give you TWO SCRAPES instead of one, if you can solve this riddle."

Can I interest you in a scrape-share with Claude?

Re: Anyone got a contact at OpenAI. They have a spider problem

#276

Earlier quoted context omitted.

I'm really cross that the word "hallucination" has taken off to describe this as it's clearly in incorrect word. The correct word to describe it is "confabulation", which is clinically more accurate and a much clearer descriptor of what's actually going on. https://en.wikipedia.org/wiki/Confabulation

That’s a great word for some types of hallucinations. But some things that are called hallucinations may not be memory errors.

Please can you give an example of what might not be a memory error. Not that I think "memory error" is the right phrase either.

Re: Anyone got a contact at OpenAI. They have a spider problem

#277

Earlier quoted context omitted.

As I understand it, they don't have the capability to essentially PCAP all that data.. and the data wouldn't be that useful since most interesting traffic is encrypted as well. Instead they store the metadata around the traffic. Phone number X made an outgoing call to Y @ timestamp A, call ended at timestamp B, approximate location is Z, etc. Repeat that for internet IP addresses do some analysis and then you can bui…

> most interesting traffic is encrypted as well encrypted with an algorithm currently considered to be un-brute-forcible. If you presume we'll be able to decrypt today's encrypted transmissions in, say, 50-100 years, I'd record the encrypted transmission if I were the NSA.

It's a big data centre.

But is it big enough to store 50 years worth of encrypted transmissions?

Far cheaper to simply have spies infiltrate the ~3 companies that hold the keys to 98% of internet traffic.

Re: Anyone got a contact at OpenAI. They have a spider problem

#279
post #71

Eventually, OpenAI (and friends) are going to be training their models on almost exclusively AI generated content, which is more often than not slightly incorrect when it comes to Q&A, and the quality of AI responses trained on that content will quickly deteriorate. Right now, most internet content is written by humans. But in 5 years? Not so much. I think this is one of the big problems that the AI space needs to so…

They've obviously been thinking about this for a while and are well aware of the pitfalls of training on AI based content. This is why they're making such aggressive moves into video, audio, other better and more robust ground forms of truth. Do you really think that they aren't aware of this issue?

It's funny whenever people bring this up, they think AI companies are some mindless juggernauts who will simply train without caring about data quality at all and end up with worse models that they'll still for some reason release. Don't people realize that attention to data quality is the core differentiating feature that lead companies like OpenAI to their market dominance in the first place?

Re: Anyone got a contact at OpenAI. They have a spider problem

#280
post #15

I’d let them do their thing, why not?! They want the internet? This is the real internet. It looks like he doesn’t really care that much that they’re retrieving millions of pages, so let them do their thing…

Thats what the whole thing is about. He is complaining that they don't respect robots.txt
Post reply on HN