Live data from Hacker News

GPT-4 details leaked?

threadreaderapp.com

611–620 of 648 posts

Re: GPT-4 details leaked?

#611
post #603

Earlier quoted context omitted.

I mean, sure you can work around it, but from your own link: >since the time and memory complexity of self-attention are quadratic in sequence length

Except in practice this is not true, and hasn't been for more than a year. It's not just a workaround either -- FlashAttention is both faster at runtime and uses less memory.

[deleted]

Re: GPT-4 details leaked?

#612

Earlier quoted context omitted.

This is incorrect: Epicycles were a model of planetary motion - the theory was that the planets moved around the earth, but also had additional circular motion as they moved along their path around the earth. This model explains the apparent geocentric motion of the planets, much as Newtonian gravity explains the apparent heliocentric motion of the planets (but is also wrong). Finding the exact parameters for the epi…

I realize I've been repeating shibboleths from my postgrad without full understanding. A problem with epicycles is they are harder to falsify. If the model doesn't match new observations, just adjust it, or add another epicycle. In contrast, Newtonian gravity can hardly be tweaked at all. So when Mercury's orbit was slightly off, they knew something was wrong. I'm not quite clear on how I feel about this. Geocentrici…

> BTW Chomsky's point E (which I'd never heard of), the last and most minor, was based on Gold's work.

Do you know where Chomsky says this exactly? I've been looking for how/if his argument is the same as Gold's.

Re: GPT-4 details leaked?

#613

Earlier quoted context omitted.

Interesting to say it's scientific because falsifiable. The objection is that it wasn't a theory, just fitting a function to data. It did "work" in that it captured some pattern: it was extremely good at generalizing/extrapolating/predicting. And was a "model" of something in the data. But there was no operational model behind it, of what was actually happening. Newtonian mechanics has a model, beyond curve fitting.…

This is incorrect: Epicycles were a model of planetary motion - the theory was that the planets moved around the earth, but also had additional circular motion as they moved along their path around the earth. This model explains the apparent geocentric motion of the planets, much as Newtonian gravity explains the apparent heliocentric motion of the planets (but is also wrong). Finding the exact parameters for the epi…

I don't think the problem wasn't that it wasn't a "model" it's that it had lots of unexplained "plug-in" behavior. The heliocentric model had much less; it explained more using less despite being less accurate.

Re: GPT-4 details leaked?

#614

Earlier quoted context omitted.

Interesting to say it's scientific because falsifiable. The objection is that it wasn't a theory, just fitting a function to data. It did "work" in that it captured some pattern: it was extremely good at generalizing/extrapolating/predicting. And was a "model" of something in the data. But there was no operational model behind it, of what was actually happening. Newtonian mechanics has a model, beyond curve fitting.…

This is incorrect: Epicycles were a model of planetary motion - the theory was that the planets moved around the earth, but also had additional circular motion as they moved along their path around the earth. This model explains the apparent geocentric motion of the planets, much as Newtonian gravity explains the apparent heliocentric motion of the planets (but is also wrong). Finding the exact parameters for the epi…

Chomsky hasn't made any reversal, Norvig misrepresents what Chomsky has said.

> Chomsky is now unhappy that LLMs can learn /any/ language, even unnatural ones, and therefore aren't good tools for understanding human language

That's also not what he has said: He said they aren't useful for understanding the human language faculty, in other words, understanding how people are able to have language. As he says it obviously can't be the same way as LLMs because LLMs are able to learn languages humans can't learn.

Re: GPT-4 details leaked?

#615

Earlier quoted context omitted.

I realize I've been repeating shibboleths from my postgrad without full understanding. A problem with epicycles is they are harder to falsify. If the model doesn't match new observations, just adjust it, or add another epicycle. In contrast, Newtonian gravity can hardly be tweaked at all. So when Mercury's orbit was slightly off, they knew something was wrong. I'm not quite clear on how I feel about this. Geocentrici…

>> RE Chomsky: You can see it's like epicycles: with enough parameters, an LLM is like a numerical method for curve fitting, that doesn't explain the data (any more than a fourier transform does). Curiously, they do seem to predict very accurately... yet also generalize strangely ("hallucinate"). What to think? Well, that's the fundamental problem of modelling: that for any set of observations there's an arbitrary nu…

> Chomsky used it to support his argument about the poverty of the stimulous but linguistics

Do you know where Chomsky refers (directly or indirectly) to Gold? I've been searching for a reference for some time.

Re: GPT-4 details leaked?

#616

Earlier quoted context omitted.

This is incorrect: Epicycles were a model of planetary motion - the theory was that the planets moved around the earth, but also had additional circular motion as they moved along their path around the earth. This model explains the apparent geocentric motion of the planets, much as Newtonian gravity explains the apparent heliocentric motion of the planets (but is also wrong). Finding the exact parameters for the epi…

I don't think the problem wasn't that it wasn't a "model" it's that it had lots of unexplained "plug-in" behavior. The heliocentric model had much less; it explained more using less despite being less accurate.

Nobody's arguing that the heliocentric model wasn't better for a variety of reasons... simply that getting superseded doesn't make the previous approach 'not science.' All models are wrong, some are useful, and all that.

And the heliocentric model still has plenty of unexplained parameters: the major and minor radii for each body. (Not to mention those pesky perturbations in mercury's orbit.)

Re: GPT-4 details leaked?

#617

Earlier quoted context omitted.

I realize I've been repeating shibboleths from my postgrad without full understanding. A problem with epicycles is they are harder to falsify. If the model doesn't match new observations, just adjust it, or add another epicycle. In contrast, Newtonian gravity can hardly be tweaked at all. So when Mercury's orbit was slightly off, they knew something was wrong. I'm not quite clear on how I feel about this. Geocentrici…

>> RE Chomsky: You can see it's like epicycles: with enough parameters, an LLM is like a numerical method for curve fitting, that doesn't explain the data (any more than a fourier transform does). Curiously, they do seem to predict very accurately... yet also generalize strangely ("hallucinate"). What to think? Well, that's the fundamental problem of modelling: that for any set of observations there's an arbitrary nu…

The gravity model is similar though: we posit a force that pulls things together, but we don't know /why/ that force seems to exist, no more than the ancients knew /why/ the planets seemed to move in smaller circles along their circular paths. We're really not /that/ enlightened, after all.

Re: GPT-4 details leaked?

#618

Earlier quoted context omitted.

It's not "a few people making less money", it's a few gigantic monoliths carving up the future, like blind watch-destroying gods -- or at least wanting to, no matter how nicely they dress it up. And it's not about utility or chance of success for everyone, either, but rather trying to do something in an ethical or more clean way just because that's more fun for them. But I have to admit to being an idealist, and whil…

Wow. Beautifully articulated. What a pleasant reply to read. I don’t have an argument regarding my position other than I agree that what you’re saying is true and that getting people like me to care and make sacrifices not just today but every day in a long term way is what makes hard problems hard.

Thank you :) But I didn't mean to say it's those pesky non-dreamers who make the problem hard. Three idealists will have ten possible utopias that are partly or completely mutually exclusive.

And to be frank, I think a lot of the finger pointing at people who don't care enough about issue X or Y is really because of not having found good ways to work constructively and make progress with however few people who do care. Partly also because people cannot agree (for long, tend to splinter into more pure sub groups and all that).

At any rate, the way can't be "I should feel bad and do better", but rather "I want what they got!". And the burden can't be on the people who aren't yet seeing anything that makes them excited to get excited anyway. And it can't be about being selfless for the benefit of others, or future generations. It has to its own reward right here and now. It is about and for you just as it is anyone else, if you know what I mean. Sacrificing others and sacrificing oneself is sacrificing people in both cases. Neither is noble IMO.

I guess the best chance of fighting tech giant strangleholds is still empowering "normal people" to carve out their own little spaces. All people will not finally learn how to make websites if only we crushed Facebook and what have you, but instead if more people had fun making their own little websites, and if we could come up with good ways for them to connect(peer-to-peer on the desktop, right after Linux!), Facebook and others would play nicer. It's not that big companies are a problem, it's the abusive things they do when they're the only game in town.

And likewise, and back on topic: a really good argument would be something I don't have the knowledge for, namely things you can do with a LLM that you can fully control (or at least can wildly poke at and experiment with, or just "download mods for") versus a much more powerful LLM that you don't really control, other than your prompts.

Thanks for reading!

Re: GPT-4 details leaked?

#619
post #79

Earlier quoted context omitted.

Have you tried falcon 40b instruct? Also take into account that chatgpt likely has some preprompt and by talking to falcon or other OS models it's all in your hands. Furthermore, Not many people discuss the significance of proper output sampling. I myself used to just test open source models with the greedy decoding only. Who knows if they wouldn't even beat (not at all)OpenAI with some clever output sampling scheme.

Idk, has anyone tried falcon yet? The support for running it remains nonexistent except for one fork of llama.cpp that isn't integrated into anything. This trend of every new model breaking compatibility really needs to stop.

I have and I am running it locally. Mostly in 7b variant which runs pretty much at chatgpt speed (when streaming) on my ryzen 3700 cpu + 32gb ram + 2xnvme in a mirror (it doesn't fit in the memory in its entirety, few gb go into the swap).

Of course to run it like that I have to be running nothing else. No xorg, no chromium etc. Just a pure linux console.

If you want to try falcon 40b instruct by yourself here I'd a public demo : https://huggingface.co/blog/falcon

Go to the bottom of the page.

Re: GPT-4 details leaked?

#620
post #606

Earlier quoted context omitted.

@Exuma, this comment is ridiculously resonant with me, the part about 'learning transaction isolation for the 50th time' is very on point too. Everything you said I pretty much feel the same way. I've accepted it as part of how I work, and the advantages are many (and valued by many) - but yes, interacting with deep experts usually ends with feeling a bit like a fraud. I feel like I maybe was an expert at whatever th…

haha yes! What mentioned about analogies... I must use like 50 analogies a day. I also noticed I can use phrases like "always" and "never" and I can say them without a second of hesitation, because they are merely indications of magnitude in a predictive sense, not a literal interpretation. But to someone who must understand information deeply, they never use phrases like that because they operate based on observed k…

Thanks everyone for this thread :)
Post reply on HN