Live data from Hacker News

GPT-4 details leaked?

threadreaderapp.com

641–648 of 648 posts

Re: GPT-4 details leaked?

#641

Earlier quoted context omitted.

I think they read the first line of your reply and took it as you arguing that it's nearly impossible for an LLM to give output that would get someone killed.

I was agreeing with them. Why assume wilful ignorance? Have we become the new reddit? People just yelling in disagreement? Maybe it's time to log out and delete this password...

Apologies, just found your post confusing and was trying to make sense.

Re: GPT-4 details leaked?

#642

Earlier quoted context omitted.

>> RE Chomsky: You can see it's like epicycles: with enough parameters, an LLM is like a numerical method for curve fitting, that doesn't explain the data (any more than a fourier transform does). Curiously, they do seem to predict very accurately... yet also generalize strangely ("hallucinate"). What to think? Well, that's the fundamental problem of modelling: that for any set of observations there's an arbitrary nu…

The gravity model is similar though: we posit a force that pulls things together, but we don't know /why/ that force seems to exist, no more than the ancients knew /why/ the planets seemed to move in smaller circles along their circular paths. We're really not /that/ enlightened, after all.

Yes, ultimately it's also descriptive.

I'd like to think explaining means giving a model simpler than the observations. But this also can be true of a purely predictive model, that offers no "why". Another commenter pointed out that epicycles do simplify - so they do "explain" in this sense.

What defines an "explanation"? What makes something a "why"?

Re: GPT-4 details leaked?

#643

Earlier quoted context omitted.

I realize I've been repeating shibboleths from my postgrad without full understanding. A problem with epicycles is they are harder to falsify. If the model doesn't match new observations, just adjust it, or add another epicycle. In contrast, Newtonian gravity can hardly be tweaked at all. So when Mercury's orbit was slightly off, they knew something was wrong. I'm not quite clear on how I feel about this. Geocentrici…

>> RE Chomsky: You can see it's like epicycles: with enough parameters, an LLM is like a numerical method for curve fitting, that doesn't explain the data (any more than a fourier transform does). Curiously, they do seem to predict very accurately... yet also generalize strangely ("hallucinate"). What to think? Well, that's the fundamental problem of modelling: that for any set of observations there's an arbitrary nu…

> not predictive models, but explanatory theories

Maybe an advantage of an explanatory theory is in revealing more of the "black box", giving more ways to check the theory. (But I'm not sure how this could apply to Newton's gravity, since the only observations were outcomes. And no plausible way to "experiment".)

> If it wasn't for the ancients stumbling and fumbling in the dark for millennia

Is there any evidence that the epicyclic models helped scientific understanding, even indirectly? Later theories didn't seem to build on it. I wonder if it actually detoured understanding, with its misleadingly impressive accuracy, so that understanding would have progressed more quickly without it.

Thinking of pg's "great work" (https://news.ycombinator.com/item?id=36550615): to be the Newton of neural nets would seem the most ambitious aspiration of our times. But it took a bunch of geniuses just to get to Newton... and it seems an even harder problem than planetary motion. Though a difference is neural nets are based on actual neurons (loosely!).

It's looking like working human-level AI will precede understanding... perhaps by those 2000 years?

Re: GPT-4 details leaked?

#644

Earlier quoted context omitted.

>> RE Chomsky: You can see it's like epicycles: with enough parameters, an LLM is like a numerical method for curve fitting, that doesn't explain the data (any more than a fourier transform does). Curiously, they do seem to predict very accurately... yet also generalize strangely ("hallucinate"). What to think? Well, that's the fundamental problem of modelling: that for any set of observations there's an arbitrary nu…

> Chomsky used it to support his argument about the poverty of the stimulous but linguistics Do you know where Chomsky refers (directly or indirectly) to Gold? I've been searching for a reference for some time.

(not who you asked) I thought this would be in the linked transcript, but it's not. Norvig must be getting it from elsewhere (maybe in the 404ed video?), but it seems like misrepresentation.

Re: GPT-4 details leaked?

#645

Earlier quoted context omitted.

I had to ask GPT what MoE means: "MoE" in the context of artificial intelligence typically stands for "Mixture of Experts". This is a machine learning technique that is based on the idea of dividing a problem into sub-problems, solving each sub-problem with a specialized "expert" (or model), and then combining their outputs.

Yep they (would) basically have 8-16 "experts" that are each about the size of GPT-3. Since they each see different batches of the dataset, they learn to model those distributions independently rather than the distribution of the whole dataset. Some of the attention is shared between them however. Then another "routing model" decides which model is most suitable for the given user prompt. Given they use relatively fe…

Please don’t post assumptions while making them look like you know 100% what you are talking about…

Re: GPT-4 details leaked?

#646

Earlier quoted context omitted.

Are you saying that this was a form of Socratic questioning- intentionally presenting an incorrect statement in order to obtain the correction?

A bit like that, but without prior knowledge of whether the information is incorrect, rather an intuition that it is correct.

Sorry but that is just ludicrous. You do not answer a factual question with a mostly made-up answer without saying clearly "this is pure speculation" at some point, preferably early on.

Re: GPT-4 details leaked?

#647
post #646

Earlier quoted context omitted.

A bit like that, but without prior knowledge of whether the information is incorrect, rather an intuition that it is correct.

Sorry but that is just ludicrous. You do not answer a factual question with a mostly made-up answer without saying clearly "this is pure speculation" at some point, preferably early on.

When I was younger (hah, I'm only 26 now) I sure did do exactly what you say. If your statement is intended to say that "People should not" then you're absolutely correct, however it is a learned behavior that some people must adopt after being corrected by their peers.

Re: GPT-4 details leaked?

#648
post #591

Earlier quoted context omitted.

The advantages in a social setting lie in the introduction of entropy, that is _creativity_, to a community. In a rigorous academic setting and with proper training these individuals are more likely identify links between ideas or information that may not seem obvious at first, and tend to be your more 'eccentric' academics. For the interests of the wider group, the best outcome is to help these individuals refine th…

You have clarified something I have always thought about very intensely and deeply but haven’t really ever read anyone else who understands that so well or rather put it into words so clearly. I’m an inferenced based learner to an extreme and it definitely has many upsides and also downsides. The upsides are being able to learn extremely rapidly by making connections between pieces of information where there’s gaps a…

I have established multiple companies, some of which have grown significantly with over 600 employees. For quite some time, I've transitioned from development and mainly held executive roles such as CEO, Chairman, etc. Simultaneously, it's intriguing to note that I've mostly been unsuccessful in securing 'normal' jobs through interviews in the past (Google, McKinsey, Bain, Accenture etc).

I believe this poses a fascinating topic on the way people assess creativity and intelligence in general.

From my perspective, the crux of the issue lies in the inherent difficulty of accurately measuring creativity in comparison to quick problem-solving skills during job interviews. Consequently, it seems that corporations tend to favor the latter.

Post reply on HN