Live data from Hacker News

Teach your LLM to answer with facts, not fiction

blog.myscale.com

81–90 of 149 posts

Re: Teach your LLM to answer with facts, not fiction

#81

'Facts' aren't as black and white as people think. "What does Charmander evolve into?" "What does the spell 'avada kedavra' do?" "What is the Sindarin word for 'friend'?" "What are the names of Santa's reindeer?" "Where did Robin Hood live?" "Where did Achilles die?" These are all 'factual questions' you can find answers to from reputable sources like Wikipedia. Google displays 'fact boxes' for several of them. Wolfr…

This is something which GPT generally isn't confused about though: it knows the answer to these questions and it knows that these are questions and statements about well-known works of fiction. I don't really think this is the source of the tendency for LLMs to make stuff up.

Re: Teach your LLM to answer with facts, not fiction

#82
post #80
post #74

Earlier quoted context omitted.

(By the way: > "protocoli(s|z)ed" the use of '-ize' is (a graecism) indicated by the OED as International English, as opposed to British, American etc. In fact, some call International English "British spelling with -ize" - it is not exactly that but close. One exception is 'analyse', but that is because linguists compromised on the "difficult" original 'analysize'.)

What's "analysize"? That's not a Greek word.

It's so determined by Fowler; pls. check this: https://www.etymonline.com/search?q=analyse

Re: Teach your LLM to answer with facts, not fiction

#83
post #82
post #80

Earlier quoted context omitted.

What's "analysize"? That's not a Greek word.

It's so determined by Fowler; pls. check this: https://www.etymonline.com/search?q=analyse

Ah, it would have been the correct way to transfer it to English, it says, not that it's in any way original.

Re: Teach your LLM to answer with facts, not fiction

#84
post #63

Earlier quoted context omitted.

> "Confabulation" And why would that be? "Hallucination" means "erratic wandering", implying one is lost - similarly to "delirium" (maetaphor using the plough) and "error". Part of the idea is that of "instead of witnessing the correct, reporting the false" - a very ancient, traditional idea, and akin to the concept of "intelligence" ( intus-legere ). "Confabulation" means locutor and interlocutor are talking, exchan…

> In psychology, confabulation is a memory error defined as the production of fabricated, distorted, or misinterpreted memories about oneself or the world. It is generally associated with certain types of brain damage (especially aneurysm in the anterior communicating artery) or a specific subset of dementias. https://en.wikipedia.org/wiki/Confabulation

Interesting, I will check that more analytically as soon as I will be back at the console,

but I am not sure - provisionally - that it can be a good idea to relate strictly human neurology to ANNs, if based on phenomena as opposed to structural issues. You do not have that problem when staying with natural language.

Re: Teach your LLM to answer with facts, not fiction

#85
post #83
post #82

Earlier quoted context omitted.

It's so determined by Fowler; pls. check this: https://www.etymonline.com/search?q=analyse

Ah, it would have been the correct way to transfer it to English, it says, not that it's in any way original.

Sorry, my imprecision. You spend ages trying to find proper expression, and yet... Well, this proves the importance of the effort.

Re: Teach your LLM to answer with facts, not fiction

#86
post #60

Earlier quoted context omitted.

Tangential: I was going to suggest "protocoli(s|z)ed" instead of protocollar, but I Googled "protocollar statements" just to check and found 2 things. First, this page was the top result! Second, "protocolar" (one ell) and "protocolary" are apparently real words. New to me, thanks.

You had me check a few sources for found expressions in use for the concept: you can find simply "protocols" (intending that), "protocol statements", the "protocol-sentence debate", "protocollar propositions"... Edit: oh, by the way, in case of interest: https://plato.stanford.edu/entries/vienna-circle/

In German they were called "Protokollsätze", which translates to "protocol sentences".

Re: Teach your LLM to answer with facts, not fiction

#87
post #85
post #83

Earlier quoted context omitted.

Ah, it would have been the correct way to transfer it to English, it says, not that it's in any way original.

Sorry, my imprecision. You spend ages trying to find proper expression, and yet... Well, this proves the importance of the effort.

Haha, that, it does.

Re: Teach your LLM to answer with facts, not fiction

#88
post #20

Earlier quoted context omitted.

As jwells89 mentioned, LLMs can extract the intention from questions and generate better queries for a search engine or database.

I don't need LLM to assume my intention, I need exact matches for keywords.

Then you don't want an LLM at all; exact keyword matches are something we did in the mid 90s, one of the specific value-adds of an LLM is that it doesn't get stuck when you've only got a half-remembered inexact quote or vague description.

Re: Teach your LLM to answer with facts, not fiction

#89
post #71

I think LLMs need to be taught to say "I don't know"/"I am not sure" or something to that effect. Another approach might be to introduce an adversarial "censor" model to guard against hallucination (or inappropriate answers).

Isn't the fundamental issue that it doesn't have any way to tell if what it thinks it knows is or isn't true? This article sounds like an idea I had independent not too long ago, but with a different goal: LLMs are great at natural language comprehension, but also have a lot of neurons dedicated to factoids. Using neurons that way is really inefficient, can we split the "language" capability from the "knowledge" capa…

I don't think that's possible in the current autocomplete based paradigm. But the LLM can be made to ignore most of its own knowledge. For example, Bing answers most questions by performing an Internet search, even if it could have answered without one. (Sometimes this makes it actually worse, e.g. when it is asked to solve a puzzle, and it gives answers to similar ones from the web, without trying to solve it itself.)

Re: Teach your LLM to answer with facts, not fiction

#90

I believe that LLMs should be banned, but if they have to exist, we should teach them ethics first before anything else.

What's your rationale for believing they should be banned? How can such algorithms be suppressed? And, is there any precedent for this, in particular that has worked?
Post reply on HN