Live data from Hacker News

GPT-4 details leaked?

threadreaderapp.com

601–610 of 648 posts

Re: GPT-4 details leaked?

#601

Earlier quoted context omitted.

This is incorrect: Epicycles were a model of planetary motion - the theory was that the planets moved around the earth, but also had additional circular motion as they moved along their path around the earth. This model explains the apparent geocentric motion of the planets, much as Newtonian gravity explains the apparent heliocentric motion of the planets (but is also wrong). Finding the exact parameters for the epi…

At some point I'd like to carefully study the history of that early era of science because I don't know it as well as I'd like. But I believe I understand that the epicyclical model (and it was a model, rather than a theory) did not in any way depend on geocentrism. For one thing, Coppernicus' model itself, while heliocentric, retained the epicycles of the earlier, geocentric model. Instead, the assumption on which t…

[deleted]

Re: GPT-4 details leaked?

#602

Earlier quoted context omitted.

Interesting to say it's scientific because falsifiable. The objection is that it wasn't a theory, just fitting a function to data. It did "work" in that it captured some pattern: it was extremely good at generalizing/extrapolating/predicting. And was a "model" of something in the data. But there was no operational model behind it, of what was actually happening. Newtonian mechanics has a model, beyond curve fitting.…

This is incorrect: Epicycles were a model of planetary motion - the theory was that the planets moved around the earth, but also had additional circular motion as they moved along their path around the earth. This model explains the apparent geocentric motion of the planets, much as Newtonian gravity explains the apparent heliocentric motion of the planets (but is also wrong). Finding the exact parameters for the epi…

I realize I've been repeating shibboleths from my postgrad without full understanding.

A problem with epicycles is they are harder to falsify. If the model doesn't match new observations, just adjust it, or add another epicycle. In contrast, Newtonian gravity can hardly be tweaked at all. So when Mercury's orbit was slightly off, they knew something was wrong.

I'm not quite clear on how I feel about this. Geocentricity is a theory, in broad terms. It seems disrespectful to say adding epicycles makes it "not a theory". As you say, there was a theory that the planets actually moved in epicyclic motion. It wasn't just calculation to them.

(I want to stress that the idea of epicycles, the mechanical craftsmanship, and actual prediction of the planets are all amazing genius.)

Yet, having more parameters than data means the model doesn't explain in simpler terms, only restates. In this sense, it's "not a theory" (by Occam's razor). It seems enough epicycles can model anything: (3Blue1Brown Fourier Series) https://youtube.com/watch?v=bL0LV0Huj1s OTOH the epicycles did predict planetary motion, so they did capture some regularity... not sure what to think.

RE Chomsky: You can see it's like epicycles: with enough parameters, an LLM is like a numerical method for curve fitting, that doesn't explain the data (any more than a fourier transform does). Curiously, they do seem to predict very accurately... yet also generalize strangely ("hallucinate"). What to think?

But it seems it's got to help! Even if only as a device, like a telescope. Also, from this interview with Terry Sejnowski (https://youtube.com/watch?v=XKC-4Tosdd8 3 hours!) there's instances where a technique was developed for a problem with Neural Nets, and an equivalent was found in the brain.

He also gives a Chomsky-like view: if you duplicated the human brain and it worked perfectly, you wouldn't have done any science if you didn't understand anything.

BTW Chomsky's point E (which I'd never heard of), the last and most minor, was based on Gold's work.

Re: GPT-4 details leaked?

#603
post #452

Earlier quoted context omitted.

During training, you have to store a dot product of Q and V that has dimension Ncrt^2. That's quadratic scaling, no?

presumably you mean a dot product of Q and K, and no you do not have to store this: https://arxiv.org/abs/2205.14135

I mean, sure you can work around it, but from your own link:

>since the time and memory complexity of self-attention are quadratic in sequence length

Re: GPT-4 details leaked?

#604
post #565
post #194

Earlier quoted context omitted.

Maybe Russia has been injecting lead into his water supply lol. I’m also saddened. Musk is not really a hero or a villain, but his manic stages have given us our first realistic shot at becoming a spacefaring civilization, and moved the needle big time on the lock that the perro cartels had on the automotive industry vis-a-vis electric cars. I hope elon gets better. Losing a billionaire tech maximalists manic episode…

It feels very ignorant to think you can diagnose someone has having manic episodes when you know nothing about them as a person, their motivations, their mental health history, and base all your opinions on mainstream outrage over tweets that are less dumb than most people’s

He has admitted it.

Re: GPT-4 details leaked?

#605
post #591

Earlier quoted context omitted.

The advantages in a social setting lie in the introduction of entropy, that is _creativity_, to a community. In a rigorous academic setting and with proper training these individuals are more likely identify links between ideas or information that may not seem obvious at first, and tend to be your more 'eccentric' academics. For the interests of the wider group, the best outcome is to help these individuals refine th…

You have clarified something I have always thought about very intensely and deeply but haven’t really ever read anyone else who understands that so well or rather put it into words so clearly. I’m an inferenced based learner to an extreme and it definitely has many upsides and also downsides. The upsides are being able to learn extremely rapidly by making connections between pieces of information where there’s gaps a…

@Exuma, this comment is ridiculously resonant with me, the part about 'learning transaction isolation for the 50th time' is very on point too.

Everything you said I pretty much feel the same way. I've accepted it as part of how I work, and the advantages are many (and valued by many) - but yes, interacting with deep experts usually ends with feeling a bit like a fraud. I feel like I maybe was an expert at whatever the thing is at some point in time, momentarily, but then I just shed the information as soon as the next thing needs to be done, and it just ends up as part of the background inference pattern matcher.

Certain things where I'm really forced to learn something deeply do stick, but I find my ways of thinking about that domain to be very different to most 'true' experts, and rely heavily on visual models and analogies with other concepts.

Re: GPT-4 details leaked?

#606
post #591

Earlier quoted context omitted.

You have clarified something I have always thought about very intensely and deeply but haven’t really ever read anyone else who understands that so well or rather put it into words so clearly. I’m an inferenced based learner to an extreme and it definitely has many upsides and also downsides. The upsides are being able to learn extremely rapidly by making connections between pieces of information where there’s gaps a…

@Exuma, this comment is ridiculously resonant with me, the part about 'learning transaction isolation for the 50th time' is very on point too. Everything you said I pretty much feel the same way. I've accepted it as part of how I work, and the advantages are many (and valued by many) - but yes, interacting with deep experts usually ends with feeling a bit like a fraud. I feel like I maybe was an expert at whatever th…

haha yes! What mentioned about analogies... I must use like 50 analogies a day. I also noticed I can use phrases like "always" and "never" and I can say them without a second of hesitation, because they are merely indications of magnitude in a predictive sense, not a literal interpretation. But to someone who must understand information deeply, they never use phrases like that because they operate based on observed knowledge and sort of "hypothesis testing" like a scientist.

It's fun to realize other people are out there who can relate. Thanks for your comment

Re: GPT-4 details leaked?

#607
post #603

Earlier quoted context omitted.

presumably you mean a dot product of Q and K, and no you do not have to store this: https://arxiv.org/abs/2205.14135

I mean, sure you can work around it, but from your own link: >since the time and memory complexity of self-attention are quadratic in sequence length

Except in practice this is not true, and hasn't been for more than a year. It's not just a workaround either -- FlashAttention is both faster at runtime and uses less memory.

Re: GPT-4 details leaked?

#608
post #578

Earlier quoted context omitted.

Yea but the bench being discussed here is FOSS. Which for me, and many, translates to can i run something useful in my closet or on my phone. I've found LLaMA neat and yea, some FOSS models are getting decent - but they're a far cry from GPT4. I pay for GPT4, use it almost daily and that's my bench. Yes, when i can run GPT4 in my closet, OpenAI will have GPT7 or w/e - but it doesn't change the fact that i have someth…

My guess is you'll be running GPT4 equivalent in your closet, but with a 4K context window. Where the big guys will have GPT-who-cares-what-version with a 100K context window. Context size is as much of a big deal as newer generations of models imo.

Am I right in my layman's understanding that context windows scaling up requires (mainly) much more compute at run time? Or do longer context models require different/longer training?

Re: GPT-4 details leaked?

#610

Earlier quoted context omitted.

This is incorrect: Epicycles were a model of planetary motion - the theory was that the planets moved around the earth, but also had additional circular motion as they moved along their path around the earth. This model explains the apparent geocentric motion of the planets, much as Newtonian gravity explains the apparent heliocentric motion of the planets (but is also wrong). Finding the exact parameters for the epi…

I realize I've been repeating shibboleths from my postgrad without full understanding. A problem with epicycles is they are harder to falsify. If the model doesn't match new observations, just adjust it, or add another epicycle. In contrast, Newtonian gravity can hardly be tweaked at all. So when Mercury's orbit was slightly off, they knew something was wrong. I'm not quite clear on how I feel about this. Geocentrici…

>> RE Chomsky: You can see it's like epicycles: with enough parameters, an LLM is like a numerical method for curve fitting, that doesn't explain the data (any more than a fourier transform does). Curiously, they do seem to predict very accurately... yet also generalize strangely ("hallucinate"). What to think?

Well, that's the fundamental problem of modelling: that for any set of observations there's an arbitrary number of models that fit the data with great accuracy and even predict future observations well; and we don't know which one is the best in the long term.

The answer is that we should prefer not predictive models, but explanatory theories, that not only predict future observations but also explain why those observations should be expected to be made.

For example, the epicyclical model did not explain anything: it said nothing about why the planets should move on circular orbits with epicycles. Kepler's laws didn't explain anything because they didn't say why the planets should move on ellpitical orbits. Newton's law of universal gravitation explained it all in one stroke: because gravity. And that's why we consider Newton the greatest scientist of his era, not Kepler, not Coppernicus, not Gallileo, but Newton, because he explained the world and didn't just describe it.

Ultimately the advantage is, like you say, that when an explanatory theory fails, we can better know why. When a predictive model fails, we have no clue.

>> BTW Chomsky's point E (which I'd never heard of), the last and most minor, was based on Gold's work.

Gold's negative learnability result was a huge upheaval that led directly to the current paradigm of machine learning. Chomsky used it to support his argument about the poverty of the stimulous but linguistics was only one of the two fields that Gold's result turned upside down.

And it was a negative result. As I say in another comment, science gives you the tools to know when you're wrong and that's how progress is made, when we find out where we were wrong before.

With epicycles, it took almost two thousand years before we figured out where the model was wrong. Let's hope that it doesn't take that long with LLMs and neural nets also, because I doubt we have another couple thousand years to spare on a wild goose chase.

>> (I want to stress that the idea of epicycles, the mechanical craftsmanship, and actual prediction of the planets are all amazing genius.)

The epicyclical model persisted for so long because it was so good, and because there was nothing better. It is common for people who don't understand science to look at scientists of the past with derision and think they weren't even scientists, but for almost two thousand years, astronomers did exactly what a scientist must do: they accepted the best available theory, even if many of them hated it with a burning passion (and they did!). If it wasn't for the ancients stumbling and fumbling in the dark for millennia, we wouldn't today be enlightened and we owe them every respect.

Post reply on HN