Live data from Hacker News

Telling GPT-4 you're scared or under pressure improves performance

aimodels.substack.com

81–90 of 255 posts

Re: Telling GPT-4 you're scared or under pressure improves performance

#81

I think this is the entry point needed to get peoples attention and explain: LLMs aren’t people, and emergent properties are being over extended. If LLMs are showing “better” performance when there are tokens that humans read as emotionally salient - Then the underlying text it’s trained on shows humans give better answers when emotionally salient context is provided. LLMs predict words. Any semantic validity is a si…

The “statistical parrot” assertion is pretty thoroughly disproven by this point, but suppose we ignore the literature and just assume it’s true: what does it matter? “Real” people are time bombs too, for instance. Is there some predictive power that we gain by reducing LLM skills to mere token production side effects?

Then why dont agents work?

If those skills were real, why do they fizzle out on production data ?

Re: Telling GPT-4 you're scared or under pressure improves performance

#82
post #72

Earlier quoted context omitted.

> The “statistical parrot” assertion is pretty thoroughly disproven by this point errr... all NNs are just optimisations of an associative probability objective: P(Y|X), they are by definition "statistical parrots". There isn't anything to prove or disprove. People offering prompts as evidence are people who fundamentally do not understand the basics. NNs aren't strange empirical objects, they're specified by mathema…

You’re confused about what “statistical parrot” means and you don’t seem to understand the difference between an optimization objective and the resulting model. The term “parrot” is used to imply inference by something akin to a look-up table, specifically it is used to indicate poor out-of-sample performance and a lack of a proper world model. The optimization objective is irrelevant when determining the generalizat…

That out of sample performance is a mirage.

Yes it’s impressive. Yes it’s got amazing zero shot performance in domains.

But there’s a pattern of failure in production which describe a limit, that shouldn’t exist if the emergent properties were stable.

You can build this right now and test it.

Build a sequence of agents to work on a domain you are not an expert in.

Let them loose. See what happens.

Do the same thing on a domain you have expertise in.

Assume the number of errors you find, the number of modifications you have to make are stable for other domains.

Re: Telling GPT-4 you're scared or under pressure improves performance

#83

Earlier quoted context omitted.

> That is why proof of concept LLM tools are mind blowing and production tools are semantic time bombs. This will be my new favorite quote for whenever someone tries to pitch his latest LLM idea

I have a whole list. Syntactic validity is not semantic validity. Word predictors not world state predictors Text prediction not fact prediction Frankly though the best answers are 1) let’s talk to infosec first 2) hey what’s the error rate ?

It's indeed a problem some people get so hyped they forget that those systems are called "language models" for a reason. They're fantastic for tasks that are as close to 100% linguistic in nature as possible, but the content might not be better than lorem ipsum in some cases, just a filler to demonstrate correct grammar.

I have noticed that GPT etc had a big "wow effect" on me, the first impression can be great because it's simply not a level of language one would expect from a computer. But prod it long enough with prompts and somehow a pattern emerges: the output never contains a higher amount of information than the prompt. Copilot can type out pages of boilerplate code because boilerplate code is noise, it can write a quicksort because the word "quicksort" already contains all the information necessary to define its behaviour.

Re: Telling GPT-4 you're scared or under pressure improves performance

#84

Earlier quoted context omitted.

The “statistical parrot” assertion is pretty thoroughly disproven by this point, but suppose we ignore the literature and just assume it’s true: what does it matter? “Real” people are time bombs too, for instance. Is there some predictive power that we gain by reducing LLM skills to mere token production side effects?

Then why dont agents work? If those skills were real, why do they fizzle out on production data ?

That's a good question. What is different about "production data"? What do those "production people" do that suddenly makes LLMs fail on things they work well on when not "production"?

Re: Telling GPT-4 you're scared or under pressure improves performance

#85
post #43

Earlier quoted context omitted.

It won't ever simulate the human brain. It may simulate human cultural knowledge, or emotions, but only as far as they are encoded in the current millennium's written knowledge. The human brain doesn't even have the concept of written language, that's all culturally learned knowledge.

How would you test whether some arbitrary thing is simulating the human brain? If you have no answer, I put it to you that your assertion is of the no-true-Scotsman type -- that is, unfalsifiable.

Its not a no-true-Scotsman type argument in a world where the existence of Scotland itself is purely hypothetical, and the arguments for the possibility of its existence in the form of a working bagpipe demonstrator are also an illustration that you don't need to be Scottish to play bagpipes.

The impossibility of testing whether the brain is adequately simulated (since a pretty basic LLM that makes no attempt to simulate the functions of a human mind yields human-like text i/o) is a point in favour of it being unlikely that we could design a fully functioning simulation of a human mind out of arbitrary material(1), since the inadequacy of testing is an impediment to actually building it.

(1) it's apparently possible for us to create human minds out of bits of humans and a process called pregnancy

Re: Telling GPT-4 you're scared or under pressure improves performance

#87
GPTs are trained on Internet data. So they are Internet simulators. If you tell the Internet you really, really need help you either get a good, quick answer or you get no answer. Because GPTs have to answer you, and they have to do it quickly, you get a good answer.

No out is talking about the many subtle trollings that have snuck into GPT training data. Nobody is wondering if the things that prompt trolling online (e.g., being rude) can cause. GPT to troll us.

Re: Telling GPT-4 you're scared or under pressure improves performance

#88

Earlier quoted context omitted.

Then why dont agents work? If those skills were real, why do they fizzle out on production data ?

That's a good question. What is different about "production data"? What do those "production people" do that suddenly makes LLMs fail on things they work well on when not "production"?

[deleted]

Re: Telling GPT-4 you're scared or under pressure improves performance

#89
post #56

Earlier quoted context omitted.

I'm curious in your statement, can you point to some papers where they addressed it?

(not op) The section A Path Forward in Managing AI Risks by Bengio et al cites a few papers: https://managing-ai-risks.com/

I read through that and none of the section (or entire work) ever talk about the above discussion. Further I looked at some of the many citations of on that section and none of them suggest that the OP is right. In fact a few of them I know disagree.
Post reply on HN