Live data from Hacker News

Truth is not a direction: a Tarski attack on LLM probes

abeljansma.nl

11–20 of 88 posts

Re: Truth is not a direction: a Tarski attack on LLM probes

#11
post #6

Aren't there definitions of Truth that are not the negation of Falsehood? Can't there be a function True(x) that is not equal to !False(x)? Can't there be a third function Paradox(x) such that these counterexamples can be considered paradoxes and therefore outside of the truth? I'm admittedly not a logician and don't formally study paradoxes, but I never quite understood the whole category of "this sentence is false"…

Interesting point about "self-referential sentences". I tend to agree. In my view a sentence saying something like "This sentence ..." does not have valid semantic meaning. It says nothing because, what "This" in "This sentence" means is ill-defined.

If terms we use are not well-defined, then sentences using such terms can not have meaning.

But for the sake of argument let's explore, what could the "this" in (so called) "self-referential" sentences refer to?

Do they refer to a specific encoding of the sentence you are reading, as some bits in computer memory perhaps?

That would require that those bits somehow have a unique "identity" and the "this" in a self-referential sentence would have to refer to those bits in specific addresses of a specific memory-chip.

But of course the "this" does not specify which memory chip, which specific (concrete) encoding of its (purported) meaning it is referring to. And if it did, then it would be talking about that specific set of bits in that specific memory-chip, not of "itself".

The fallacy is that what we perceive as a "sentence we read" is somehow "speaking" of something. But no, the sentence is not a subject, a sentence can not speak, and THEREFORE it specifically can not speak of itself.

A sentence can not speak, only actors, only subjects, like humans and AI, can "speak". And their speech must be encoded in some physical medium. A written sentence like "This sentence is ..." gives us the false impressions that somehow the SENTENCE IS SPEAKING of itself!

But speech can not speak, speech is the product of speaking.

Hence, in my view, "self-referential sentences" do not have any meaning and whatever paradoxes they might seem to create are results of confusion between ontological levels of "Subject" vs. "Speech".

Re: Truth is not a direction: a Tarski attack on LLM probes

#12
post #6

Aren't there definitions of Truth that are not the negation of Falsehood? Can't there be a function True(x) that is not equal to !False(x)? Can't there be a third function Paradox(x) such that these counterexamples can be considered paradoxes and therefore outside of the truth? I'm admittedly not a logician and don't formally study paradoxes, but I never quite understood the whole category of "this sentence is false"…

There is no lack of more elaborate logics:

https://en.wikipedia.org/wiki/Non-classical_logic

Re: Truth is not a direction: a Tarski attack on LLM probes

#13
post #6

Aren't there definitions of Truth that are not the negation of Falsehood? Can't there be a function True(x) that is not equal to !False(x)? Can't there be a third function Paradox(x) such that these counterexamples can be considered paradoxes and therefore outside of the truth? I'm admittedly not a logician and don't formally study paradoxes, but I never quite understood the whole category of "this sentence is false"…

Fuzzy logic?

Re: Truth is not a direction: a Tarski attack on LLM probes

#14
post #6

Aren't there definitions of Truth that are not the negation of Falsehood? Can't there be a function True(x) that is not equal to !False(x)? Can't there be a third function Paradox(x) such that these counterexamples can be considered paradoxes and therefore outside of the truth? I'm admittedly not a logician and don't formally study paradoxes, but I never quite understood the whole category of "this sentence is false"…

I’m not a logician either but believe this is what Tarski’s definition of truth solves for. In order to make a statement about the statement itself, you have to introduce a new meta language. Then a statement in the meta language is only true if the underlying statement is true.

Much more rigorous explanation: https://plato.stanford.edu/entries/tarski-truth/

Re: Truth is not a direction: a Tarski attack on LLM probes

#15
post #7

Earlier quoted context omitted.

If anything, the whole vector space is the LLM's truth.

LLMs are not optimized only for truth they are optimized for a more complicated objective that includes e.g. humans liking their output. It is a universal truth that to get humans to like you, you have to lie to them.

> It is a universal truth that to get humans to like you, you have to lie to them.

Thus said HAL9000

Re: Truth is not a direction: a Tarski attack on LLM probes

#16
post #8

A direction that is 99.99% accurate survives this argument completely. For all practical purposes one does not need totality.

A basic course in statistics will inform you of why a 99.99% accurate test should be looked at with skepticism when diagnosing a rare disease. Yet we see the fancy 9s and think somehow this many 9s is enough.

Re: Truth is not a direction: a Tarski attack on LLM probes

#17
post #16
post #8

A direction that is 99.99% accurate survives this argument completely. For all practical purposes one does not need totality.

A basic course in statistics will inform you of why a 99.99% accurate test should be looked at with skepticism when diagnosing a rare disease. Yet we see the fancy 9s and think somehow this many 9s is enough.

Sure, but that’s not what we are talking about

Re: Truth is not a direction: a Tarski attack on LLM probes

#18

I think this article pushes the premise farther than is reasonable. The best anyone expects from an LLM "truth vector" is that it would encode the model's belief about whether the statement is true. Of course a perfect truth oracle is impossible.

> The best anyone expects from an LLM "truth vector" is that it would encode the model's belief about whether the statement is true.

I think even that's too-optimistic: The LLM is a document-extender, so its "belief" is whether a token seems like it would statistically fit-next in a partial document, based on prior documents. This is usually not the kind of analytic truth we're interested in, and we've already figured out how to constantly extract it.

If we peek at vectors and weights, we'll we'll probably end up measuring the moods and styles for whatever tokens are about to get emitted next, whether that's dialogue for a fictional character (of various kinds), a narrator, or an impersonal memo conclusion paragraph. We'll be measuring "earnestness and conviction", on the same level as "loquaciousness" or "pleading" or "talking like a pirate."

So is Truthiness [0] what we really want? Probably not. If our document described the character as Yoda, then The Force connecting all existence ends up truthy. Using "a really gullible person" can repeat anything you supply as truthy. Even if we set things up as "a respected encyclopedia article" or "a relentlessly logical super-genius", we're really changing the influence mix of styles and biases, rather than creating a logical mind independent of text inside the LLM.

[0] https://en.wikipedia.org/wiki/Truthiness

Re: Truth is not a direction: a Tarski attack on LLM probes

#20
post #6

Aren't there definitions of Truth that are not the negation of Falsehood? Can't there be a function True(x) that is not equal to !False(x)? Can't there be a third function Paradox(x) such that these counterexamples can be considered paradoxes and therefore outside of the truth? I'm admittedly not a logician and don't formally study paradoxes, but I never quite understood the whole category of "this sentence is false"…

Interesting point about "self-referential sentences". I tend to agree. In my view a sentence saying something like "This sentence ..." does not have valid semantic meaning. It says nothing because, what "This" in "This sentence" means is ill-defined. If terms we use are not well-defined, then sentences using such terms can not have meaning. But for the sake of argument let's explore, what could the "this" in (so call…

As the article notes, the sentence "This sentence is written in English" is well understood, and true. "This sentence is written in French" is also well understood and false.
Post reply on HN