Live data from Hacker News

Anyone got a contact at OpenAI. They have a spider problem

mailman.nanog.org

291–300 of 400 posts

Re: Anyone got a contact at OpenAI. They have a spider problem

#292
post #86

Earlier quoted context omitted.

I wonder how much the source content is the cause of hallucinations rather than anything inherent to LLMs. I mean if someone posts a question on an internet forum that I don't know the answer to, I'm certainly not going to post "I don't know" since that wouldn't be useful. In fact, in general, in any non one-on-one conversation the answer "I don't know" is not useful because if you don't know in a group, your silence…

Reminds me of a joke Three logicians walk into a bar. The bartender says "what'll it be, three beers?" The first logician says "I don't know". The second logician says "I don't know". The third logician says "Yes".

If, like me, you didn't get the joke at first:

Both of the first two logicians wanted a beer; otherwise they would know the answer was "no". The third logician recognizes this, and therefore knows the answer.

Re: Anyone got a contact at OpenAI. They have a spider problem

#293

Earlier quoted context omitted.

This must be why self-driving cars always ignore the speed limit. ;)

More directly, e.g. Tesla boasts of training their FSD on data captured from their customer's unassisted driving. So it's hardly surprising that it imitates a lot of humans' bad habits, e.g. rolling past stop lines.

AI DRIVR claims that beta V12 is much better precisely because it takes rules less literally and drives more naturally.

Re: Anyone got a contact at OpenAI. They have a spider problem

#295
post #141

Earlier quoted context omitted.

This sounds interesting, could you give more detail on what you're referring to?

I'm referring to his two works, the "Tractatus Logico-Philosophicus" and "Philosophical Investigations". There's a lot explored here, but Wittgenstein basically makes the argument that the natural logic of language—how we deduce meaning from terms in a context and naturally disambiguate the semantics of ambiguous phrases—is different from the sort of formal propositional logic that forms the basis of western philosop…

What we cannot speak about we must pass over in silence.

Re: Anyone got a contact at OpenAI. They have a spider problem

#296
post #141

Earlier quoted context omitted.

This sounds interesting, could you give more detail on what you're referring to?

I'm referring to his two works, the "Tractatus Logico-Philosophicus" and "Philosophical Investigations". There's a lot explored here, but Wittgenstein basically makes the argument that the natural logic of language—how we deduce meaning from terms in a context and naturally disambiguate the semantics of ambiguous phrases—is different from the sort of formal propositional logic that forms the basis of western philosop…

> some people seem to have wildly different thresholds for the propositional coherence they can spot

This sums up the last decade remarkably well.

Re: Anyone got a contact at OpenAI. They have a spider problem

#297
post #95

Earlier quoted context omitted.

That’s a good observation. If LLMs had taken off 15 years ago, maybe they would answer every question with “this has already been asked before. Please use the search function”

Marked as duplicate.

Damn. I was hoping to be the fastest gun in the west on this one!

Re: Anyone got a contact at OpenAI. They have a spider problem

#298

Earlier quoted context omitted.

Any data scraped would be instantly deduplicated after the fact by whatever semantic dedupe engine they've cooked up.

What has it got to do with deduplication? I'm talking about crafting some kind of alternative (not necessarily duplicate) data. I agree some kind of post data collection cleaning/filtering of the data before training could potentially catch it. But maybe not!

Ah fair enough. The OP here mentioned having highly similar content on each of the many domains.

Re: Anyone got a contact at OpenAI. They have a spider problem

#299
post #71

Eventually, OpenAI (and friends) are going to be training their models on almost exclusively AI generated content, which is more often than not slightly incorrect when it comes to Q&A, and the quality of AI responses trained on that content will quickly deteriorate. Right now, most internet content is written by humans. But in 5 years? Not so much. I think this is one of the big problems that the AI space needs to so…

It is already solved. Look at how Microsoft trained Phi - they used existing models to generate synthetic data from textbooks. That allowed them to create a new dataset grounded in “fact” at a far higher quality than common crawl or others. It looks less like an ouroboros and more like a bootstrapping problem.

Is this like, the AI equivalent of “another layer will fix it” that crypto fans used?

“It’s ok bro, another model will fix, just please, one more ~layer~ ~agent~ model”

It’s all fun and games until you can’t reliably generate your base models anymore, because all your _base_ data is too polluted.

Let’s not forget MS has a $10bn stake in the current crop of LLM’s turning out to be as magic as they claim, so I’m sure they will do anything to ensure that happens.

Re: Anyone got a contact at OpenAI. They have a spider problem

#300

Earlier quoted context omitted.

Reminds me of a joke Three logicians walk into a bar. The bartender says "what'll it be, three beers?" The first logician says "I don't know". The second logician says "I don't know". The third logician says "Yes".

If, like me, you didn't get the joke at first: Both of the first two logicians wanted a beer; otherwise they would know the answer was "no". The third logician recognizes this, and therefore knows the answer.

Unless one of those wanted two beers. Or 0.5 beer. Or -1 beers. Or 1e9 beers. Or 2147483648 beers.
Post reply on HN