We added them to our robots.txt, but traffic didn't stop. I complained to their provider, who happened to also be AWS. Oh, the shock - they didn't care and recommended using robots.txt! AWS was making good money on both ends with this bot that apparently had more money to burn than we do.
Anyone got a contact at OpenAI. They have a spider problem
311–320 of 400 posts
Re: Anyone got a contact at OpenAI. They have a spider problem
#312I'm so tired of bots. A certain bot from Singapore started pulling all of the product images across multiple domains. Ok, whatever... Then we realized it never stopped. It made enough requests to download them 4-5x over and was still going. The AWS bill was not nice. We added them to our robots.txt, but traffic didn't stop. I complained to their provider, who happened to also be AWS. Oh, the shock - they didn't care…
Re: Anyone got a contact at OpenAI. They have a spider problem
#313Earlier quoted context omitted.
Part of me really wants to believe this is a joke. But given how often toddlers are "taken care of" by planting them in front of youtube :|
Which is crazy because there's plenty of good content for kids on Youtube (if you really need a break!). Blippy, Meekah, Seasame Street, even that mind-numbing drivel Cocomelon (which at least got my girls talking/singing really early).
Re: Anyone got a contact at OpenAI. They have a spider problem
#314Earlier quoted context omitted.
You didn't really do the joke right. The second person is supposed to be doing the thing requested, but in a way the first person doesn't like. In this case, that would be AI companies contributing to a "free and open internet", but doing it """wrong""". But they're not contributing at all. The problem isn't that "free and open internet" is getting monkeys-pawed, it's all this closed proprietary stuff.
> You didn't really do the joke right. It summed up many comments and attitudes I see here on LLM's in under 30 words. In fact it's so clever it's one of the few posts here I'm certain wasn't produced by an LLM.
It's almost clever.
Re: Anyone got a contact at OpenAI. They have a spider problem
#315I'm so tired of bots. A certain bot from Singapore started pulling all of the product images across multiple domains. Ok, whatever... Then we realized it never stopped. It made enough requests to download them 4-5x over and was still going. The AWS bill was not nice. We added them to our robots.txt, but traffic didn't stop. I complained to their provider, who happened to also be AWS. Oh, the shock - they didn't care…
Re: Anyone got a contact at OpenAI. They have a spider problem
#316Isn’t it funny that all the “worthless” content out there on the internet is actually changing the world. Like how 4chan was mocked as being a cesspit for losers, but now everyone knows memes like Pepe the frog and Wojack from there. And like now this very comment and the billions of other comments on here, Reddit, Twitter, etc that are regarded as a “waste of time” are being used to train multi billion dollar compan…
It's possible to hold those two thoughts in mind at the same time.
Re: Anyone got a contact at OpenAI. They have a spider problem
#317Re: Anyone got a contact at OpenAI. They have a spider problem
#318Earlier quoted context omitted.
He says 3 million, and 1.8 million are for robots.txt So 1.2 million non robots.txt requests, when his robots.txt file is configured as follows # buzz off User-agent: GPTBot Disallow: / Theoretically if they were actually respecting robots.txt they wouldn't crawl any pages on the site. Which would also mean they wouldn't be following any links... aka not finding the N subdomains.
His site has a subdomain for every page, and the crawler is considering those each to be unique sites.
Re: Anyone got a contact at OpenAI. They have a spider problem
#319Earlier quoted context omitted.
I'm really cross that the word "hallucination" has taken off to describe this as it's clearly in incorrect word. The correct word to describe it is "confabulation", which is clinically more accurate and a much clearer descriptor of what's actually going on. https://en.wikipedia.org/wiki/Confabulation
I proposed[1] the portmanteau "hallucofabulation" as a compromise, but it hasn't caught on yet. I'm totally shocked and dismayed by this, of course. [1]: https://news.ycombinator.com/item?id=36977935
Re: Anyone got a contact at OpenAI. They have a spider problem
#320Earlier quoted context omitted.
I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…
>In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence indicates that. This isn't true. There are many contexts where it is true but it doesn't actually generalize they way you say it does. There are plenty of cases where experts in a non-one-on-one context will express a lack of knowledge. Sometimes this will be as part of m…