Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

311–320 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#311
I'm so tired of bots. A certain bot from Singapore started pulling all of the product images across multiple domains. Ok, whatever... Then we realized it never stopped. It made enough requests to download them 4-5x over and was still going. The AWS bill was not nice.

We added them to our robots.txt, but traffic didn't stop. I complained to their provider, who happened to also be AWS. Oh, the shock - they didn't care and recommended using robots.txt! AWS was making good money on both ends with this bot that apparently had more money to burn than we do.

Re: Anyone got a contact at OpenAI. They have a spider problem

#312

I'm so tired of bots. A certain bot from Singapore started pulling all of the product images across multiple domains. Ok, whatever... Then we realized it never stopped. It made enough requests to download them 4-5x over and was still going. The AWS bill was not nice. We added them to our robots.txt, but traffic didn't stop. I complained to their provider, who happened to also be AWS. Oh, the shock - they didn't care…

This is exactly why we need to move past “the Web” to precisely pay for storage and hosting with micropayments. It is the requester who should pay!

Re: Anyone got a contact at OpenAI. They have a spider problem

#313

Earlier quoted context omitted.

Part of me really wants to believe this is a joke. But given how often toddlers are "taken care of" by planting them in front of youtube :|

Which is crazy because there's plenty of good content for kids on Youtube (if you really need a break!). Blippy, Meekah, Seasame Street, even that mind-numbing drivel Cocomelon (which at least got my girls talking/singing really early).

There's actually no such thing as good "content" for kids, sorry.

Re: Anyone got a contact at OpenAI. They have a spider problem

#314

Earlier quoted context omitted.

You didn't really do the joke right. The second person is supposed to be doing the thing requested, but in a way the first person doesn't like. In this case, that would be AI companies contributing to a "free and open internet", but doing it """wrong""". But they're not contributing at all. The problem isn't that "free and open internet" is getting monkeys-pawed, it's all this closed proprietary stuff.

> You didn't really do the joke right. It summed up many comments and attitudes I see here on LLM's in under 30 words. In fact it's so clever it's one of the few posts here I'm certain wasn't produced by an LLM.

The factual elements work as a summary. But the entire aspect of just deserts, getting what you asked for and being unhappy, hypocrisy, it isn't there. So by implying that kind of thing, it doesn't work.

It's almost clever.

Re: Anyone got a contact at OpenAI. They have a spider problem

#315

I'm so tired of bots. A certain bot from Singapore started pulling all of the product images across multiple domains. Ok, whatever... Then we realized it never stopped. It made enough requests to download them 4-5x over and was still going. The AWS bill was not nice. We added them to our robots.txt, but traffic didn't stop. I complained to their provider, who happened to also be AWS. Oh, the shock - they didn't care…

Use Cloudflare? Just give their server a captcha challenge?

Re: Anyone got a contact at OpenAI. They have a spider problem

#316

Isn’t it funny that all the “worthless” content out there on the internet is actually changing the world. Like how 4chan was mocked as being a cesspit for losers, but now everyone knows memes like Pepe the frog and Wojack from there. And like now this very comment and the billions of other comments on here, Reddit, Twitter, etc that are regarded as a “waste of time” are being used to train multi billion dollar compan…

> Like how 4chan was mocked as being a cesspit for losers, but now everyone knows memes like Pepe the frog and Wojack from there.

It's possible to hold those two thoughts in mind at the same time.

Re: Anyone got a contact at OpenAI. They have a spider problem

#317

Earlier quoted context omitted.

And the human bartender passes the check to the third logician.

The third logician never finishes his beer, his friends get more free beers. The bar overflows.

How much would that exploit be worth on the open market?

Re: Anyone got a contact at OpenAI. They have a spider problem

#318
post #65

Earlier quoted context omitted.

He says 3 million, and 1.8 million are for robots.txt So 1.2 million non robots.txt requests, when his robots.txt file is configured as follows # buzz off User-agent: GPTBot Disallow: / Theoretically if they were actually respecting robots.txt they wouldn't crawl any pages on the site. Which would also mean they wouldn't be following any links... aka not finding the N subdomains.

His site has a subdomain for every page, and the crawler is considering those each to be unique sites.

Of course it’s considering them as unique sites. They are unique sites.

Re: Anyone got a contact at OpenAI. They have a spider problem

#319

Earlier quoted context omitted.

I'm really cross that the word "hallucination" has taken off to describe this as it's clearly in incorrect word. The correct word to describe it is "confabulation", which is clinically more accurate and a much clearer descriptor of what's actually going on. https://en.wikipedia.org/wiki/Confabulation

I proposed[1] the portmanteau "hallucofabulation" as a compromise, but it hasn't caught on yet. I'm totally shocked and dismayed by this, of course. [1]: https://news.ycombinator.com/item?id=36977935

It would probably help if you didn't drop the "n" from the confabulation part.

Re: Anyone got a contact at OpenAI. They have a spider problem

#320
post #119
post #86

Earlier quoted context omitted.

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…

>In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence indicates that. This isn't true. There are many contexts where it is true but it doesn't actually generalize they way you say it does. There are plenty of cases where experts in a non-one-on-one context will express a lack of knowledge. Sometimes this will be as part of m…

I personally will almost always say I don't know while talking thru to a solution. Admittedly this is informal speech that doesn't make it to written form.
Post reply on HN