Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

331–340 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#331
post #150

I'm more interested in what that content farm is for. It looks pointless, but I suspect there's a bizarre economic incentive. There are affiliate links, but how much could that possibly bring in?

It's for shits-and-giggles and it's doing its job really well right now. Not everything needs to have an economic purpose, 100 trackers, ads and backed by a company.

And yet it has the Amazon links, which makes it appear to have some economic purpose...

Re: Anyone got a contact at OpenAI. They have a spider problem

#332

Earlier quoted context omitted.

I'm not sure any publisher means for their robots.txt to be read as: "You're disallowed, but go head and slurp the content anyway so you can look for external links or any indication that maybe you are allowed to digest this material anyway, and then interpret that how you'd like. I trust you to know what's best and I'm sure you kind of get the gist of what I mean here."

How would one know he is disallowed without reading each site?

The convention is that crawlers first read /robots.txt to see what they're encouraged to scrape and what they're not meant to, and then hopefully honor those directions.

In this case, as in many, the disallow rules are intentionally meant to protect the signal quality and efficiency of the crawler.

Re: Anyone got a contact at OpenAI. They have a spider problem

#333

Earlier quoted context omitted.

Please can you give an example of what might not be a memory error. Not that I think "memory error" is the right phrase either.

I was thinking along the lines of answering with correct information but not following the prompts. Maybe this could be considered confabulation also.

My view is that hallucination is something related to the interpretation of reality. It's not really directly mapping to memory at all. The mechanisms of confabulation entirely surround the gluing together of memories, and what are these models other than some sort of representation of memory.

I believe that you can also cause something a bit like a transient dysphasia by giving them bad inputs as well, so there is that on the language production side. However there's still nothing that pertains to the experience aspects central to what hallucinations actually are.

Re: Anyone got a contact at OpenAI. They have a spider problem

#334

Earlier quoted context omitted.

Which is crazy because there's plenty of good content for kids on Youtube (if you really need a break!). Blippy, Meekah, Seasame Street, even that mind-numbing drivel Cocomelon (which at least got my girls talking/singing really early).

There's actually no such thing as good "content" for kids, sorry.

If you are trying to get kids to be fluent in a 2nd or 3rd language, there certainly is

Re: Anyone got a contact at OpenAI. They have a spider problem

#335

Earlier quoted context omitted.

I'm really cross that the word "hallucination" has taken off to describe this as it's clearly in incorrect word. The correct word to describe it is "confabulation", which is clinically more accurate and a much clearer descriptor of what's actually going on. https://en.wikipedia.org/wiki/Confabulation

I think hallucination is better and more accurate at it implies a bit of imagination and buffoonery deceit. I don’t think confabulate matches as well as it implies confusion or mixture of different ideas. ChatGPT isn’t confused, it’s making things up. It’s trying to bullshit as best it can in hope that what it makes up convinces its user.

That making things up based on memories of past things is entirely what confabulation is. Bullshitting in the large as it were. I've met quite a few clinical confabulators (people with Korsakov syndrome and the like) and I find the parallels remarkable.

Re: Anyone got a contact at OpenAI. They have a spider problem

#336

Earlier quoted context omitted.

Reminds me of a joke Three logicians walk into a bar. The bartender says "what'll it be, three beers?" The first logician says "I don't know". The second logician says "I don't know". The third logician says "Yes".

That sounds like my old neighbour, a professor of logic from the university of science.

I've heard tell of the place. By chance, did he have a doghouse?

Re: Anyone got a contact at OpenAI. They have a spider problem

#338
post #294
post #266

Am I the only one who was hoping—even though I knew it wouldn’t be the case—that OpenAI’s server farm was infested with actual spiders and they were getting into other people’s racks?

very xkcd

Going way back to the start

Re: Anyone got a contact at OpenAI. They have a spider problem

#339

Earlier quoted context omitted.

Reminds me of a joke Three logicians walk into a bar. The bartender says "what'll it be, three beers?" The first logician says "I don't know". The second logician says "I don't know". The third logician says "Yes".

If, like me, you didn't get the joke at first: Both of the first two logicians wanted a beer; otherwise they would know the answer was "no". The third logician recognizes this, and therefore knows the answer.

why is it three logicians? wouldn't it work with just two?

Re: Anyone got a contact at OpenAI. They have a spider problem

#340

Earlier quoted context omitted.

Is this like, the AI equivalent of “another layer will fix it” that crypto fans used? “It’s ok bro, another model will fix, just please, one more ~layer~ ~agent~ model” It’s all fun and games until you can’t reliably generate your base models anymore, because all your _base_ data is too polluted. Let’s not forget MS has a $10bn stake in the current crop of LLM’s turning out to be as magic as they claim, so I’m sure t…

I mean, you can use Phi now and it outclasses anything else in its size. This isn’t some “it could happen” situation.

Oh I’m sure it works wonderfully for now.

My point is about the inevitable future when _those_ models start to struggle.

The phi approach doesn’t seem like breaking the ouroboros, it just feels like inserting another model/snake into the loop.

Post reply on HN