Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

351–360 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#351

Earlier quoted context omitted.

Part of me really wants to believe this is a joke. But given how often toddlers are "taken care of" by planting them in front of youtube :|

Which is crazy because there's plenty of good content for kids on Youtube (if you really need a break!). Blippy, Meekah, Seasame Street, even that mind-numbing drivel Cocomelon (which at least got my girls talking/singing really early).

Sure, and most of it starts with a jingle and ends up with begging block.

I used to cut all those things to shape with youtube-dl and Audacity; we have a library of a good hundred+ of sanitized songs to play, but with modern world hating files and anything offline, it turned out to be quite a hassle to keep the practice up.

Re: Anyone got a contact at OpenAI. They have a spider problem

#352
post #86

Earlier quoted context omitted.

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…

I'm really cross that the word "hallucination" has taken off to describe this as it's clearly in incorrect word. The correct word to describe it is "confabulation", which is clinically more accurate and a much clearer descriptor of what's actually going on. https://en.wikipedia.org/wiki/Confabulation

It's "fabrication", plain and simple.

Fully agreed that "hallucination" is a bonkers word for it — sensational and melodramatic. But few people know what a confabulation is, and moreover it's an overly complex way to describe the phenomenon.

The LLM is making something up. It's a fabrication.

It's not fanciful; it's not spooky; it's mundane, as it should be.

Re: Anyone got a contact at OpenAI. They have a spider problem

#353
post #345

Earlier quoted context omitted.

I get the sentiment, but when reality hits unrealistic parental expectations, things get messy. If you have to put a show on TV to give some songs to sing along to or to distract them while you're making lunch, I'm not judging you, and I think it's best to put this content on a gradient rather than black and white.

For all of human history until 70 years ago, no baby watched TV. Reconsider what you "have to" do.

People had extended family around, now days they don't.

Parenting is easier if you have 5 family members in walking distance and they also have similarly ages kids who can all play together.

Re: Anyone got a contact at OpenAI. They have a spider problem

#354
post #55

Earlier quoted context omitted.

Except the first thing openai does is read robots.txt. However, robots.txt doesn't cover multiple domains, and every link that's being crawled is to a new domain, which requires a new read of a robots .txt on the new domain.

And they do have (the same) robots.txt on every domain, tailored for GPTbot, i.e. https://petra-cody-carlene.web.sp.am/robots.txt So, GPTBot is not following robots.txt, apparently.

[deleted]

Re: Anyone got a contact at OpenAI. They have a spider problem

#355
post #185

Earlier quoted context omitted.

That's an explanation of why their answers can be useful, but doesn't relate to their ability to "not know" an answer

I suspect this is going to be a disagreement on the meaning of "to know". On the same lines as why people argue if a tree falling in a wood where nobody can hear it makes sound because some people implicitly regard sound is the qualia while others regard it as the vibrations in the air.

LLMs don't know anything except the most frequent observed response to a context made up of a sequence of tokens.

How often do the words "I don't know" get uttered in books, papers, articles, stack overflow, or any other resource of knowledge?

Re: Anyone got a contact at OpenAI. They have a spider problem

#356

Earlier quoted context omitted.

The re-use of the "c" as a soft c in "hallucinate" and then a hard c in confabulate is confusing, and probably affecting the uptake of your neologism.

Maybe if I added a hyphen? "halluco-fabulation"?

would lead to hafabulation. Still not acceptable in my view - confabulation is clearly the correct term

Re: Anyone got a contact at OpenAI. They have a spider problem

#357

Earlier quoted context omitted.

I'm really cross that the word "hallucination" has taken off to describe this as it's clearly in incorrect word. The correct word to describe it is "confabulation", which is clinically more accurate and a much clearer descriptor of what's actually going on. https://en.wikipedia.org/wiki/Confabulation

It's "fabrication", plain and simple. Fully agreed that "hallucination" is a bonkers word for it — sensational and melodramatic. But few people know what a confabulation is, and moreover it's an overly complex way to describe the phenomenon. The LLM is making something up. It's a fabrication. It's not fanciful; it's not spooky; it's mundane, as it should be.

Fabrication implies making things up for the sake of it. Confabulation is similar but is defined by making things up due to some limitation in capacity.

Re: Anyone got a contact at OpenAI. They have a spider problem

#358
post #55

Earlier quoted context omitted.

It's a honeypot. He's telling people openai doesn't respect robots.txt and just scrapes whatever the hell it wants.

Except the first thing openai does is read robots.txt. However, robots.txt doesn't cover multiple domains, and every link that's being crawled is to a new domain, which requires a new read of a robots .txt on the new domain.

[deleted]

Re: Anyone got a contact at OpenAI. They have a spider problem

#359

Earlier quoted context omitted.

I'm really cross that the word "hallucination" has taken off to describe this as it's clearly in incorrect word. The correct word to describe it is "confabulation", which is clinically more accurate and a much clearer descriptor of what's actually going on. https://en.wikipedia.org/wiki/Confabulation

I proposed[1] the portmanteau "hallucofabulation" as a compromise, but it hasn't caught on yet. I'm totally shocked and dismayed by this, of course. [1]: https://news.ycombinator.com/item?id=36977935

but confabulation sufficiently describes the phenomenon without the need for ugly neologisms.

Re: Anyone got a contact at OpenAI. They have a spider problem

#360
Anyone care to explain the purpose of Levine's https://www.web.sp.am site. Are the names randomly generated. Pardon my ignorance.

This is the type of stuff the news organisations should be publishing about "AI". Instead I keep reading or hearing people referring to training data with phrases like, "The sum of all human knowledge..." Quite shocking anyone would believe that.

Post reply on HN