Earlier quoted context omitted.
You can make it invent a new language: https://maximumeffort.substack.com/p/i-taught-chatgpt-to-inv... I am sure you will continue to argue that this is still in line with everything-thats-ever-written prediction but my opinion is that at that point, it's a meaningless distinction. The human brain is also just a machine.
The brain is a machine, the issue is the difference between 2 claims LLMs are enough to be a brain LLMs are not enough to be a brain.
Telling GPT-4 you're scared or under pressure improves performance
151–160 of 255 posts
Re: Telling GPT-4 you're scared or under pressure improves performance
#152Earlier quoted context omitted.
Cameras don't have eyeballs, Microphones don't have hair cells, Speakers don't have vocal cords, processors don't don't do arithmetic with neurons, yet we all agree that they are capable of emulating the meaningful aspects of these functions. All of your claims are either incorrect (not adapting, expressing desires, beliefs, preferences, ...) or fail to eliminate irrelevant differences. If we're to have any sensible…
> an implementation detail Yip, so I deny this premise. I take it to be the heart of the matter. > we might as well throw away 80% of our current scientific understanding Yip, i'd be down for that. Though maybe i'd say, 30-40%. Science in the strongest sense has no theory-building need for statistics. Those areas of science which have only statistical models, and not causal-ontological ones aren't science -- and i'd…
With Plato's cave, the scientists do not put literally every possible object in front of the light, they sample the shadow representation and, again, construct a model around those samples.
Also, you're describing statistics in an incredibly dismissive way. Stats is decidedly not just "taking the average". At the very least not in this brainless, first-order way you describe here.
Let's explore this with an thought experiment:
A model of some process has been confirmed across the globe, at least 5000 studies show the same result. Yet, one day, a study is published that fails to demonstrate the desired effect. Without using statistics, please tell me which action should be taken next:
A: The stray result is investigated for experimental failures.
B: The entire model of the process is immediately dismissed and we start from scratch.
By the way, you're welcome to call me uninformed, but I'd ask you to at least provide either your credentials or research that directly contradicts me.
Oh, I almost forgot. I know all of these definitions, please actually engage with what I'm saying instead of insinuating that I'm missing information.
Re: Telling GPT-4 you're scared or under pressure improves performance
#153Earlier quoted context omitted.
You can make it invent a new language: https://maximumeffort.substack.com/p/i-taught-chatgpt-to-inv... I am sure you will continue to argue that this is still in line with everything-thats-ever-written prediction but my opinion is that at that point, it's a meaningless distinction. The human brain is also just a machine.
So I was with a financial researcher recently, and he wanted to use ChatGPT to summarise some reference financial data -- and it did so, actually correctly. Being sceptical, as every person ought in these matters, I changed the finical data and performed the same analysis (both in a new tab, and within the same convo). The results were the same! How strange? Well, in being reference financial data ChatGPT was reporti…
It's not, though. It is in fact able to summarize financial data, just as it's able to write code and diagnose a medical condition. It makes mistakes, yes, even grave ones, much more so than experts in those fields would.
Re: Telling GPT-4 you're scared or under pressure improves performance
#154Earlier quoted context omitted.
I'd phrase characterizing the reliability of out-of-sample performance a priori as impossible, but not necessarily automatically failing. There may be a subtle correlation between properties needed to answer a specific out-of-sample request and in-sample features. Unfortunately, prior to training/testing and without recognizing that correlation in the data set, I believe it's impossible to guarantee the model will in…
In essence: “You cant know in advance how far the model can approximate semantic patterns” So claiming that out-of-sample performance is a mirage, would be a bridge too far?
Re: Telling GPT-4 you're scared or under pressure improves performance
#155Earlier quoted context omitted.
No, my first paragraph sidesteps messy debate about what intelligence is and focuses on the simple fact a human brain simulation exists purely in the realms of the hypothetical. Its not the no true Scotsman fallacy unless the Scotsman actually exists.
What is the "working bagpipe demonstrator" in your analogy? And what is the meaning of "you don't need to be Scottish"?
So it's literally the inverse of "no true Scotsman". We don't have Scotsmen not doing things that all "true Scotsmen" are supposed to do, we have "definitely not Scotsmen" passing benchmarks for Scottishness (in a world in which Scotland itself is only an aspiration)
Re: Telling GPT-4 you're scared or under pressure improves performance
#156Earlier quoted context omitted.
> an implementation detail Yip, so I deny this premise. I take it to be the heart of the matter. > we might as well throw away 80% of our current scientific understanding Yip, i'd be down for that. Though maybe i'd say, 30-40%. Science in the strongest sense has no theory-building need for statistics. Those areas of science which have only statistical models, and not causal-ontological ones aren't science -- and i'd…
The scientific method is inherently statistical, we take a finite amount of observations and construct a model that best represents those observations. So yes, sorry, I should have said 100%. With Plato's cave, the scientists do not put literally every possible object in front of the light, they sample the shadow representation and, again, construct a model around those samples. Also, you're describing statistics in…
The model F=GMm/r^2, for example, has a causal and ontological semantics: F is a force, M a mass etc. these are pieces of reality. And this formula (though actual a little suspicious in many ways, GR fixes this) nevertheless says there is a force between masses that has certain properties etc.
Now you can say that astrologers who recorded positions of the stars in books helped 'create' this model in the sense that this data was inspiration to newton. But he didnt derive the model from this data: there are an infinite number of (causal) models consistent with the data (statistical models).
Rather newton played around with creating geometries, just like the vase-makers in plato's cave. Newton built various ways the world might be first, projected data out of them, and compared that to 'the statistical data of his day' (ie., astrology).
There's nothing in the data to tell Newton he was right. Indeed, vast amounts of it told him it was wrong: such a law does not describe the known solar system at his time, very far away from it.
Nevertheless 'modelling shadows' isnt science; and his job was science. So one has to compare actual explanatory models, and his was the best.
What you're describing above is hypothesis testing which occurs long after theory building. Broader theories create causal models, causal models create sets of predictions, we call some subset a hypothesis and by hypothesis testing we can select, in an often psuedoscientific way, between causal models.
This technique occurs long after the invention of science, arises out of explanatorily bankrupt areas, as a way of 'giving researchers something to do'. It's wholly pointless without theory-building, it is just averaging shadows.
The science we think of when using the term 'Science' owes very very little to the modern practice of hypothesis testing. Comparing hypotheses is an intellectual part of assessing explanations -- identifiable formal statistical methods entered in the early 20th C.
For almost all of scientific history 'data' functions much more like reductio-ad-absurdum premises in philosophical arguments than as sets of numbers from which to derive distributions.
That latter system, in most cases, fails. It provies a wholly illusory sense that data can decide matters; and applies in cases requiring extreme non-physical assumptions (eg., of the normalcy of the underlying data, or of a fast rate of convergence of the central limit theorem).
Much real-world phenomena studied by stats cannot really be studied by data analysis at all; and the whole method of 20th C. statistical hypothesis testing is the opening sales pitch to entire fields of pseudoscience.
Re: Telling GPT-4 you're scared or under pressure improves performance
#157Earlier quoted context omitted.
"Are not" is the rub here. They 100% are not those things... but they also approximate them well-enough to be functionally useful. I.e. the high-dimensional curve-fitting / compression conceptualization of ML, which intuitively expresses both its strengths and weaknesses. If "it" is represented in the data set (explicitly or implicitly), the "curve" will fit to that property. Simultaneously, the "curve" is approximat…
They're statistical approximations of these things -- that's really the rub. You can approximate a human capacity, say theory-of-mind, with another kind of ape: play some hide-and-seek game. You can approximate the knowledge of a trivia-master with a child and a trivia book. These are quite different sorts of approximations. A 1/100th scale bridge build to stand for a real one is quite different than taking some prio…
I'd only add that certainty requirements for real-world target applications can differ substantially. I.e. engineering vs art.
A toy box that gives magic answers 85% of the time is incredibly useful in some scenarios -- e.g. seeding the beginning of a manual research process with initial topics.
Which seems a general rule of thumb for deploying GenAI into production these days: find use cases where there are minimal consequences to being occasionally wrong.
In my head, that's the Netflix recommendation test. What's the impact if Netflix gives me a bad recommendation? Consequently, that's how aggressive they can be with their models.
Re: Telling GPT-4 you're scared or under pressure improves performance
#158Earlier quoted context omitted.
So I was with a financial researcher recently, and he wanted to use ChatGPT to summarise some reference financial data -- and it did so, actually correctly. Being sceptical, as every person ought in these matters, I changed the finical data and performed the same analysis (both in a new tab, and within the same convo). The results were the same! How strange? Well, in being reference financial data ChatGPT was reporti…
>Since it's incapable of actually summarising financial data It's not, though. It is in fact able to summarize financial data, just as it's able to write code and diagnose a medical condition. It makes mistakes, yes, even grave ones, much more so than experts in those fields would.
Do you see a difference between the process of adding numbers and dividing by their count (taking a mean) and emitting numeric tokens which are most probable for a given input?
The former is called "taking a mean" the latter isnt. This system never engages in any method to summarise financial data. It's method is always the same: to emit tokens most probable given a set of historical tokens.
It's the difference between saying "the average of 1,2,3" is 2 because that sentence occurs 1,000,000 times and saying it's 2 because you've literally computed it.
This system does not run financial summary algorithms. It's a trick
Re: Telling GPT-4 you're scared or under pressure improves performance
#159Re: Telling GPT-4 you're scared or under pressure improves performance
#160Earlier quoted context omitted.
They're statistical approximations of these things -- that's really the rub. You can approximate a human capacity, say theory-of-mind, with another kind of ape: play some hide-and-seek game. You can approximate the knowledge of a trivia-master with a child and a trivia book. These are quite different sorts of approximations. A 1/100th scale bridge build to stand for a real one is quite different than taking some prio…
We're in agreement on the nature of SOTA and the world, I think. I'd only add that certainty requirements for real-world target applications can differ substantially. I.e. engineering vs art. A toy box that gives magic answers 85% of the time is incredibly useful in some scenarios -- e.g. seeding the beginning of a manual research process with initial topics. Which seems a general rule of thumb for deploying GenAI in…
You can't predict almost anything anyway. The relevant distribution for acting is a utility+risk distribution over various sets of predictions.
If you compute that, most of ML/AI isnt very useful for most of anyone. It's kinda interesting that generative AI finally achieved something here, given our 'distributions for action'.
Of course the popular conversation isnt there yet, people still think these things work. When more people find their utility hit by their failures, i think we'll arrive back to status-quo-ante-openai where people turn off siri.
Nevertheless -- there is an achivement here (finicial in paying for training; and legal in whitewashing copyright away) -- which will have some impact