Earlier quoted context omitted.
And they do have (the same) robots.txt on every domain, tailored for GPTbot, i.e. https://petra-cody-carlene.web.sp.am/robots.txt So, GPTBot is not following robots.txt, apparently.
All the lines related to GPTBot are commented out. That robots.txt isn't trying to block it. Either it has been changed recently or most of this comment thread is mistaken.
Anyone got a contact at OpenAI. They have a spider problem
271–280 of 400 posts
Re: Anyone got a contact at OpenAI. They have a spider problem
#272Honeypots like this seem like a super interesting way to poison LLM training.
Any data scraped would be instantly deduplicated after the fact by whatever semantic dedupe engine they've cooked up.
Re: Anyone got a contact at OpenAI. They have a spider problem
#273Re: Anyone got a contact at OpenAI. They have a spider problem
#274Earlier quoted context omitted.
Yeah Microsoft, I mean openai, compressing it into an ”ai voice” is “open”.
Hackernews really instant able to understand dark jokes are they.
Re: Anyone got a contact at OpenAI. They have a spider problem
#275Earlier quoted context omitted.
Here you go (1 req/min, 10 bytes/sec), please report results :) http { limit_req_zone $binary_remote_addr zone=ten_bytes_per_second:10m rate=1r/m; server { location / { if ($http_user_agent = "mimo") { limit_req zone=ten_bytes_per_second burst=5; limit_rate 10; } } } }
Scrapers of the future won't be ifElse logic, they will be LLM agents themselves. The slow loris robots.txt has to provide an interface to it's own LLM, which engages the scraper LLM in conversation, aiming to extend it as long as possible. "OK I will tell you whether or not I can be scraped. BUT FIRST, listen to this offer. I can give you TWO SCRAPES instead of one, if you can solve this riddle."
Re: Anyone got a contact at OpenAI. They have a spider problem
#276Earlier quoted context omitted.
I'm really cross that the word "hallucination" has taken off to describe this as it's clearly in incorrect word. The correct word to describe it is "confabulation", which is clinically more accurate and a much clearer descriptor of what's actually going on. https://en.wikipedia.org/wiki/Confabulation
That’s a great word for some types of hallucinations. But some things that are called hallucinations may not be memory errors.
Re: Anyone got a contact at OpenAI. They have a spider problem
#277Earlier quoted context omitted.
As I understand it, they don't have the capability to essentially PCAP all that data.. and the data wouldn't be that useful since most interesting traffic is encrypted as well. Instead they store the metadata around the traffic. Phone number X made an outgoing call to Y @ timestamp A, call ended at timestamp B, approximate location is Z, etc. Repeat that for internet IP addresses do some analysis and then you can bui…
> most interesting traffic is encrypted as well encrypted with an algorithm currently considered to be un-brute-forcible. If you presume we'll be able to decrypt today's encrypted transmissions in, say, 50-100 years, I'd record the encrypted transmission if I were the NSA.
But is it big enough to store 50 years worth of encrypted transmissions?
Far cheaper to simply have spies infiltrate the ~3 companies that hold the keys to 98% of internet traffic.
Re: Anyone got a contact at OpenAI. They have a spider problem
#278Re: Anyone got a contact at OpenAI. They have a spider problem
#279Eventually, OpenAI (and friends) are going to be training their models on almost exclusively AI generated content, which is more often than not slightly incorrect when it comes to Q&A, and the quality of AI responses trained on that content will quickly deteriorate. Right now, most internet content is written by humans. But in 5 years? Not so much. I think this is one of the big problems that the AI space needs to so…
It's funny whenever people bring this up, they think AI companies are some mindless juggernauts who will simply train without caring about data quality at all and end up with worse models that they'll still for some reason release. Don't people realize that attention to data quality is the core differentiating feature that lead companies like OpenAI to their market dominance in the first place?
Re: Anyone got a contact at OpenAI. They have a spider problem
#280I’d let them do their thing, why not?! They want the internet? This is the real internet. It looks like he doesn’t really care that much that they’re retrieving millions of pages, so let them do their thing…