Live data from Hacker News

Truth is not a direction: a Tarski attack on LLM probes

abeljansma.nl

31–40 of 88 posts

Re: Truth is not a direction: a Tarski attack on LLM probes

#31
post #8

A direction that is 99.99% accurate survives this argument completely. For all practical purposes one does not need totality.

For a one time coin flip, sure. For a lot of them, depending on the stakes, that can be very, very wrong.

https://www.google.com/search?q=why+99.99+accuracy+is+not+en...

It doesn't take a lot of thought experiments to realize this. Imagine if every bite of food we eat had a 0.01% chance to turn into something instantly lethal in our mouth. Average lifespans would be reduced measurably, and apart from anxiety, we'd develop all sorts of strategies and laws around that. E.g. absolutely NO eating for airplane pilots. You wouldn't go on a date to have dinner, dancing and sex, you'd go dancing and have sex, and then have breakfast. People would modify their jaws and stomachs so they could eat less, but bigger chunks of food. It would be a whole thing!

And that's not even talking about water changing on us, or a tiny chance of getting sucked into the toilet whenever we use it, and a lot of other things where going from damn near 100% to 99.99% would change everything for the worse, by so much.

Re: Truth is not a direction: a Tarski attack on LLM probes

#32
post #19

I wish I could get a model to state its assumptions.

Models do not have assumptions. They have probabilities for what the next token should be. With enough context, in the context window and built into the model, that next token isn't completely random, it's correlated with something someone might choose to write.

But people write all kinds of crap manually. So far, the data people have been writing has tended to be denser around what people could agree on (there are many lies, but only one truth), so the model is more likely to go there.

If we start putting AI generated text into the training data, it's not clear what that means for the resulting model. It's already clear that some actors are trying to influence models by putting large amounts of content out there that agree with them.

Figuring out which content is safe to train from is the real problem for future model trainers.

Re: Truth is not a direction: a Tarski attack on LLM probes

#33
post #17
post #16

Earlier quoted context omitted.

A basic course in statistics will inform you of why a 99.99% accurate test should be looked at with skepticism when diagnosing a rare disease. Yet we see the fancy 9s and think somehow this many 9s is enough.

Sure, but that’s not what we are talking about

Diagnosis of rare diseases is a subset of "true" statements, no?

Re: Truth is not a direction: a Tarski attack on LLM probes

#34
post #16

Earlier quoted context omitted.

A basic course in statistics will inform you of why a 99.99% accurate test should be looked at with skepticism when diagnosing a rare disease. Yet we see the fancy 9s and think somehow this many 9s is enough.

Really? Laws of probability works against your argument. Such a basic course in statistics needs to be scrutinized.

Every breath you take, there is a 99.99% chance that everything is normal, and a 0.01% chance that you breathe mild acid which horribly burns and causes a massive coughing fit.

It probably don’t cause long term damage unless that breath happened to be more important than normal, like while driving right as a child runs into the road.

Would you act differently knowing that you had one of these occasional acid breaths? Even if it only happened once per day on average (0.001%)

Re: Truth is not a direction: a Tarski attack on LLM probes

#35
I think that one of the main problems with LLMs is that we're shoveling everything in to them, without any guidance as to what is "real"/"true"/"factual" - crazy anti-vaxxer-cooker stuff is is sitting in their along with science without any concept of scientific reality and no guidance to what is real and what is insane minds spinning on each other

Re: Truth is not a direction: a Tarski attack on LLM probes

#36

> It might seem absurd to you to even suggest superhuman AIs could function as a truth-oracle (it certainly does to me), but there are two reasons to take it seriously. First, it is how these things will be used practically by the vast majority of people. They are already replacing standard Google search results, and I’ve had many discussions end with people delegating final authority on the truth to an AI. There are…

We can call it AI Engine Optimization and it will become alright...

Re: Truth is not a direction: a Tarski attack on LLM probes

#37

> It might seem absurd to you to even suggest superhuman AIs could function as a truth-oracle (it certainly does to me), but there are two reasons to take it seriously. First, it is how these things will be used practically by the vast majority of people. They are already replacing standard Google search results, and I’ve had many discussions end with people delegating final authority on the truth to an AI. There are…

> people delegating final authority on the truth to an AI

I can't avoid looking down on this attitude. But speaks more about people than it speaks about AI: those people want to win an argument, nothing more and nothing else

> There are already organizations taking advantage of this, actively producing "AI propaganda", meaning propaganda aimed at the LLMs themselves in order to influence their understanding of what is truthful and bend it towards powerful actors' agendas. They're not even hiding the fact that they're doing this.

namely grok.

Re: Truth is not a direction: a Tarski attack on LLM probes

#38

> It might seem absurd to you to even suggest superhuman AIs could function as a truth-oracle (it certainly does to me), but there are two reasons to take it seriously. First, it is how these things will be used practically by the vast majority of people. They are already replacing standard Google search results, and I’ve had many discussions end with people delegating final authority on the truth to an AI. There are…

> The set-up assumes that the game and life are the same thing, and such is the pervasive nature of the idea of the game within the society that just by believing that, they make it so.

(The player of Games, banks)

Re: Truth is not a direction: a Tarski attack on LLM probes

#39

I think this article pushes the premise farther than is reasonable. The best anyone expects from an LLM "truth vector" is that it would encode the model's belief about whether the statement is true. Of course a perfect truth oracle is impossible.

Not only that, but nobody really cares about whether truth vectors can work correctly under carefully constructed paradox edge cases. Well, maybe mathematicians do, but nobody else.

Re: Truth is not a direction: a Tarski attack on LLM probes

#40
post #28

Title is a bit clickbaitish, but the content is well worth reading - came in with my pitchfork ready and left agreeing with basically all of it, with questions like ‘what if the probe could return 3 dimensions: truthfulness, knowledge confidence and decidability?’ Also the observation that people treat LLMs like oracles when they’re everything but is spot on, something I’ve also been thinking about and it’s quite a b…

> what if the probe could return 3 dimensions: truthfulness, knowledge confidence and decidability?

can't be done because the other two dimensions depend on the first. or in other words it's all the same dimension only with a different name.

Post reply on HN