Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

341–350 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#341

Earlier quoted context omitted.

If, like me, you didn't get the joke at first: Both of the first two logicians wanted a beer; otherwise they would know the answer was "no". The third logician recognizes this, and therefore knows the answer.

why is it three logicians? wouldn't it work with just two?

[Three is always funnier](https://en.wikipedia.org/wiki/Rule_of_three_(writing)).

Re: Anyone got a contact at OpenAI. They have a spider problem

#342

I'm so tired of bots. A certain bot from Singapore started pulling all of the product images across multiple domains. Ok, whatever... Then we realized it never stopped. It made enough requests to download them 4-5x over and was still going. The AWS bill was not nice. We added them to our robots.txt, but traffic didn't stop. I complained to their provider, who happened to also be AWS. Oh, the shock - they didn't care…

Use Cloudflare? Just give their server a captcha challenge?

Block their origin network...

Re: Anyone got a contact at OpenAI. They have a spider problem

#343

Earlier quoted context omitted.

If, like me, you didn't get the joke at first: Both of the first two logicians wanted a beer; otherwise they would know the answer was "no". The third logician recognizes this, and therefore knows the answer.

why is it three logicians? wouldn't it work with just two?

I recently heard this explained () in the following way: three is the smallest number where you can set up an expectation (with the first two) and then break it. This is why three is such a common number, not just in jokes but in all sorts of story-telling.

() In a lecture by the mathematician & author Sarah Hart.

Re: Anyone got a contact at OpenAI. They have a spider problem

#344
post #139

Earlier quoted context omitted.

Nowadays they are instead learning to say "please join our Discord for support"!

Much like one of the first phrases spoken by babies today is "like and subscribe".

Google was depressingly early on my daughter's word list (from, "hey Google, play $FOO")

Re: Anyone got a contact at OpenAI. They have a spider problem

#345

Earlier quoted context omitted.

There's actually no such thing as good "content" for kids, sorry.

I get the sentiment, but when reality hits unrealistic parental expectations, things get messy. If you have to put a show on TV to give some songs to sing along to or to distract them while you're making lunch, I'm not judging you, and I think it's best to put this content on a gradient rather than black and white.

For all of human history until 70 years ago, no baby watched TV. Reconsider what you "have to" do.

Re: Anyone got a contact at OpenAI. They have a spider problem

#346
post #86

Earlier quoted context omitted.

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…

I'm really cross that the word "hallucination" has taken off to describe this as it's clearly in incorrect word. The correct word to describe it is "confabulation", which is clinically more accurate and a much clearer descriptor of what's actually going on. https://en.wikipedia.org/wiki/Confabulation

[deleted]

Re: Anyone got a contact at OpenAI. They have a spider problem

#347
post #266

Am I the only one who was hoping—even though I knew it wouldn’t be the case—that OpenAI’s server farm was infested with actual spiders and they were getting into other people’s racks?

I was hoping that some large set of keywords generated spider images.

Dear Jane,

Can I have my drawing of a spider back then please?

Re: Anyone got a contact at OpenAI. They have a spider problem

#348
post #86

Earlier quoted context omitted.

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…

I've wondered if one could train a LLM on a closed set of curated knowledge. Then include training data that models the behaviour of not knowing. To the point that it could generalize to being able to represent its own not knowing. Because expecting a behaviour, like knowing you don't know, that isn't represented in the training set is silly. Kids make stuff up at first, then we correct them - so they have a way to l…

No it won’t work.

This is not a brain. The best analogy is an English major.

They are good at language, not reasoning.

Humans see language and think reason. It seems we can’t separate the two.

Re: Anyone got a contact at OpenAI. They have a spider problem

#349

I'm so tired of bots. A certain bot from Singapore started pulling all of the product images across multiple domains. Ok, whatever... Then we realized it never stopped. It made enough requests to download them 4-5x over and was still going. The AWS bill was not nice. We added them to our robots.txt, but traffic didn't stop. I complained to their provider, who happened to also be AWS. Oh, the shock - they didn't care…

I've been there.

Can't you just 404 the bot in your reverse proxy, optionally setting logging to /dev/null?

Still annoying, but a smaller AWS bill. And maybe the requester will eventually get the point.

Re: Anyone got a contact at OpenAI. They have a spider problem

#350

Earlier quoted context omitted.

Much like one of the first phrases spoken by babies today is "like and subscribe".

Part of me really wants to believe this is a joke. But given how often toddlers are "taken care of" by planting them in front of youtube :|

It's not. Kids overhear what parents watch, their ears are like little recorders. Meanwhile, kids videos near-universally end with either a like&subscribe admonition, or some crap like "ask parents to download our tablet app". Even the quality videos, they all do that.

Even if you don't show children videos, but want to play some music, YouTube is still the least-hassle, least-bullshit music stream player (arguably still it's main use for adults, too). Ain't anyone got time to deal with Spotify's ever more broken app. And this is the limit of technical skill of almost all parents. They can't exactly run SponsorBlock in YouTube's mobile app (and paid YouTube doesn't help here either, surprise surprise).

Not making excuses (though I'm not really blaming parents for this) - just saying how things actually are.

Post reply on HN