Live data from Hacker News

Telling GPT-4 you're scared or under pressure improves performance

aimodels.substack.com

151–160 of 255 posts

Re: Telling GPT-4 you're scared or under pressure improves performance

#151

Earlier quoted context omitted.

You can make it invent a new language: https://maximumeffort.substack.com/p/i-taught-chatgpt-to-inv... I am sure you will continue to argue that this is still in line with everything-thats-ever-written prediction but my opinion is that at that point, it's a meaningless distinction. The human brain is also just a machine.

The brain is a machine, the issue is the difference between 2 claims LLMs are enough to be a brain LLMs are not enough to be a brain.

[deleted]

Re: Telling GPT-4 you're scared or under pressure improves performance

#152

Earlier quoted context omitted.

Cameras don't have eyeballs, Microphones don't have hair cells, Speakers don't have vocal cords, processors don't don't do arithmetic with neurons, yet we all agree that they are capable of emulating the meaningful aspects of these functions. All of your claims are either incorrect (not adapting, expressing desires, beliefs, preferences, ...) or fail to eliminate irrelevant differences. If we're to have any sensible…

> an implementation detail Yip, so I deny this premise. I take it to be the heart of the matter. > we might as well throw away 80% of our current scientific understanding Yip, i'd be down for that. Though maybe i'd say, 30-40%. Science in the strongest sense has no theory-building need for statistics. Those areas of science which have only statistical models, and not causal-ontological ones aren't science -- and i'd…

The scientific method is inherently statistical, we take a finite amount of observations and construct a model that best represents those observations. So yes, sorry, I should have said 100%.

With Plato's cave, the scientists do not put literally every possible object in front of the light, they sample the shadow representation and, again, construct a model around those samples.

Also, you're describing statistics in an incredibly dismissive way. Stats is decidedly not just "taking the average". At the very least not in this brainless, first-order way you describe here.

Let's explore this with an thought experiment:

A model of some process has been confirmed across the globe, at least 5000 studies show the same result. Yet, one day, a study is published that fails to demonstrate the desired effect. Without using statistics, please tell me which action should be taken next:

A: The stray result is investigated for experimental failures.

B: The entire model of the process is immediately dismissed and we start from scratch.

By the way, you're welcome to call me uninformed, but I'd ask you to at least provide either your credentials or research that directly contradicts me.

Oh, I almost forgot. I know all of these definitions, please actually engage with what I'm saying instead of insinuating that I'm missing information.

Re: Telling GPT-4 you're scared or under pressure improves performance

#153

Earlier quoted context omitted.

You can make it invent a new language: https://maximumeffort.substack.com/p/i-taught-chatgpt-to-inv... I am sure you will continue to argue that this is still in line with everything-thats-ever-written prediction but my opinion is that at that point, it's a meaningless distinction. The human brain is also just a machine.

So I was with a financial researcher recently, and he wanted to use ChatGPT to summarise some reference financial data -- and it did so, actually correctly. Being sceptical, as every person ought in these matters, I changed the finical data and performed the same analysis (both in a new tab, and within the same convo). The results were the same! How strange? Well, in being reference financial data ChatGPT was reporti…

>Since it's incapable of actually summarising financial data

It's not, though. It is in fact able to summarize financial data, just as it's able to write code and diagnose a medical condition. It makes mistakes, yes, even grave ones, much more so than experts in those fields would.

Re: Telling GPT-4 you're scared or under pressure improves performance

#154
post #135

Earlier quoted context omitted.

I'd phrase characterizing the reliability of out-of-sample performance a priori as impossible, but not necessarily automatically failing. There may be a subtle correlation between properties needed to answer a specific out-of-sample request and in-sample features. Unfortunately, prior to training/testing and without recognizing that correlation in the data set, I believe it's impossible to guarantee the model will in…

In essence: “You cant know in advance how far the model can approximate semantic patterns” So claiming that out-of-sample performance is a mirage, would be a bridge too far?

Maybe "a mirage that might actually be true"? Which is a terrible thing to rely on! Unless it's usually true?

Re: Telling GPT-4 you're scared or under pressure improves performance

#155

Earlier quoted context omitted.

No, my first paragraph sidesteps messy debate about what intelligence is and focuses on the simple fact a human brain simulation exists purely in the realms of the hypothetical. Its not the no true Scotsman fallacy unless the Scotsman actually exists.

What is the "working bagpipe demonstrator" in your analogy? And what is the meaning of "you don't need to be Scottish"?

An LLM. Not a brain simulation, but mimics the text outputs of brain activity quite well just by parsing text. Turns out that like getting an Englishman to play bagpipes, you can get an LLM to write as if it's angry or drunk or horny (A corollary of a sufficiently large text learning model generating convincingly angry outputs based on pure word association is that you can't trust an attempt to build a long running process with adrenaline and cortisol analogues has adequately simulated emotional state just because the communications module writes convincingly angry responses)

So it's literally the inverse of "no true Scotsman". We don't have Scotsmen not doing things that all "true Scotsmen" are supposed to do, we have "definitely not Scotsmen" passing benchmarks for Scottishness (in a world in which Scotland itself is only an aspiration)

Re: Telling GPT-4 you're scared or under pressure improves performance

#156

Earlier quoted context omitted.

> an implementation detail Yip, so I deny this premise. I take it to be the heart of the matter. > we might as well throw away 80% of our current scientific understanding Yip, i'd be down for that. Though maybe i'd say, 30-40%. Science in the strongest sense has no theory-building need for statistics. Those areas of science which have only statistical models, and not causal-ontological ones aren't science -- and i'd…

The scientific method is inherently statistical, we take a finite amount of observations and construct a model that best represents those observations. So yes, sorry, I should have said 100%. With Plato's cave, the scientists do not put literally every possible object in front of the light, they sample the shadow representation and, again, construct a model around those samples. Also, you're describing statistics in…

Well if you think scientific models are associative statistical models there is some information missing in your view, I'd say. Since, well, they arent.

The model F=GMm/r^2, for example, has a causal and ontological semantics: F is a force, M a mass etc. these are pieces of reality. And this formula (though actual a little suspicious in many ways, GR fixes this) nevertheless says there is a force between masses that has certain properties etc.

Now you can say that astrologers who recorded positions of the stars in books helped 'create' this model in the sense that this data was inspiration to newton. But he didnt derive the model from this data: there are an infinite number of (causal) models consistent with the data (statistical models).

Rather newton played around with creating geometries, just like the vase-makers in plato's cave. Newton built various ways the world might be first, projected data out of them, and compared that to 'the statistical data of his day' (ie., astrology).

There's nothing in the data to tell Newton he was right. Indeed, vast amounts of it told him it was wrong: such a law does not describe the known solar system at his time, very far away from it.

Nevertheless 'modelling shadows' isnt science; and his job was science. So one has to compare actual explanatory models, and his was the best.

What you're describing above is hypothesis testing which occurs long after theory building. Broader theories create causal models, causal models create sets of predictions, we call some subset a hypothesis and by hypothesis testing we can select, in an often psuedoscientific way, between causal models.

This technique occurs long after the invention of science, arises out of explanatorily bankrupt areas, as a way of 'giving researchers something to do'. It's wholly pointless without theory-building, it is just averaging shadows.

The science we think of when using the term 'Science' owes very very little to the modern practice of hypothesis testing. Comparing hypotheses is an intellectual part of assessing explanations -- identifiable formal statistical methods entered in the early 20th C.

For almost all of scientific history 'data' functions much more like reductio-ad-absurdum premises in philosophical arguments than as sets of numbers from which to derive distributions.

That latter system, in most cases, fails. It provies a wholly illusory sense that data can decide matters; and applies in cases requiring extreme non-physical assumptions (eg., of the normalcy of the underlying data, or of a fast rate of convergence of the central limit theorem).

Much real-world phenomena studied by stats cannot really be studied by data analysis at all; and the whole method of 20th C. statistical hypothesis testing is the opening sales pitch to entire fields of pseudoscience.

Re: Telling GPT-4 you're scared or under pressure improves performance

#157
post #122

Earlier quoted context omitted.

"Are not" is the rub here. They 100% are not those things... but they also approximate them well-enough to be functionally useful. I.e. the high-dimensional curve-fitting / compression conceptualization of ML, which intuitively expresses both its strengths and weaknesses. If "it" is represented in the data set (explicitly or implicitly), the "curve" will fit to that property. Simultaneously, the "curve" is approximat…

They're statistical approximations of these things -- that's really the rub. You can approximate a human capacity, say theory-of-mind, with another kind of ape: play some hide-and-seek game. You can approximate the knowledge of a trivia-master with a child and a trivia book. These are quite different sorts of approximations. A 1/100th scale bridge build to stand for a real one is quite different than taking some prio…

We're in agreement on the nature of SOTA and the world, I think.

I'd only add that certainty requirements for real-world target applications can differ substantially. I.e. engineering vs art.

A toy box that gives magic answers 85% of the time is incredibly useful in some scenarios -- e.g. seeding the beginning of a manual research process with initial topics.

Which seems a general rule of thumb for deploying GenAI into production these days: find use cases where there are minimal consequences to being occasionally wrong.

In my head, that's the Netflix recommendation test. What's the impact if Netflix gives me a bad recommendation? Consequently, that's how aggressive they can be with their models.

Re: Telling GPT-4 you're scared or under pressure improves performance

#158

Earlier quoted context omitted.

So I was with a financial researcher recently, and he wanted to use ChatGPT to summarise some reference financial data -- and it did so, actually correctly. Being sceptical, as every person ought in these matters, I changed the finical data and performed the same analysis (both in a new tab, and within the same convo). The results were the same! How strange? Well, in being reference financial data ChatGPT was reporti…

>Since it's incapable of actually summarising financial data It's not, though. It is in fact able to summarize financial data, just as it's able to write code and diagnose a medical condition. It makes mistakes, yes, even grave ones, much more so than experts in those fields would.

It isnt making mistakes ... its never actually doing it.

Do you see a difference between the process of adding numbers and dividing by their count (taking a mean) and emitting numeric tokens which are most probable for a given input?

The former is called "taking a mean" the latter isnt. This system never engages in any method to summarise financial data. It's method is always the same: to emit tokens most probable given a set of historical tokens.

It's the difference between saying "the average of 1,2,3" is 2 because that sentence occurs 1,000,000 times and saying it's 2 because you've literally computed it.

This system does not run financial summary algorithms. It's a trick

Re: Telling GPT-4 you're scared or under pressure improves performance

#159
So you have a training set of questions and answers generated by humans. For simplification you ask ten people the same question and get 10 answers and feed it to llm and then in testing time ask the same question. Now say there is a correct answer. Now in training set you asked the question normally 5 times and added “this is very important” 5 times. And it turns out humans have better answers in the training set if you added the qualifier. And during testing time, when you add the qualifier you are telling the llm to put more weight on those 5 answers, and it performs better. Just like in image generation you need lots of negative prompts because the training set is so dirty.

Re: Telling GPT-4 you're scared or under pressure improves performance

#160
post #157

Earlier quoted context omitted.

They're statistical approximations of these things -- that's really the rub. You can approximate a human capacity, say theory-of-mind, with another kind of ape: play some hide-and-seek game. You can approximate the knowledge of a trivia-master with a child and a trivia book. These are quite different sorts of approximations. A 1/100th scale bridge build to stand for a real one is quite different than taking some prio…

We're in agreement on the nature of SOTA and the world, I think. I'd only add that certainty requirements for real-world target applications can differ substantially. I.e. engineering vs art. A toy box that gives magic answers 85% of the time is incredibly useful in some scenarios -- e.g. seeding the beginning of a manual research process with initial topics. Which seems a general rule of thumb for deploying GenAI in…

That's exactly my advise also. Take a risk profile over the predictions, and use them only when the risk of error is low.

You can't predict almost anything anyway. The relevant distribution for acting is a utility+risk distribution over various sets of predictions.

If you compute that, most of ML/AI isnt very useful for most of anyone. It's kinda interesting that generative AI finally achieved something here, given our 'distributions for action'.

Of course the popular conversation isnt there yet, people still think these things work. When more people find their utility hit by their failures, i think we'll arrive back to status-quo-ante-openai where people turn off siri.

Nevertheless -- there is an achivement here (finicial in paying for training; and legal in whitewashing copyright away) -- which will have some impact

Post reply on HN