Live data from Hacker News

Extracting concepts from GPT-4

openai.com

71–80 of 155 posts

Re: Extracting concepts from GPT-4

#71

This is super cool, it feels like going in the direction of the "deep"/high level type of semantic searching I've been waiting for. I like their examples of basically filtering documents for the "concept" of price increases, or even something as high level as a rhetorical question I wonder how this compares to training/fine tuning a model on examples of rhetorical questions and asking it to find it in a given documen…

Exa is trying to do this. I've found some sort of interesting stuff this way but it honestly doesn't feel quite good enough yet to me.

https://exa.ai/search?c=all

Re: Extracting concepts from GPT-4

#72

Earlier quoted context omitted.

I read this as "we have not built up tools / math to understand neural networks as they are new and exciting" and not as "neural networks are magical and complex and not understandable because we are meddling with something we cannot control". A good example would be planes - it took a long while to develop mathematical models that could be used to model behavior. Meanwhile practical experimentation developed decent…

Chaotic nonlinear dynamics have been an object of mathematical research for a very long time and we have built up good mathematical tools to work with them, but in spite of that turbulent flow and similar phenomena (brains/LLM's) remain poorly understood. The problem is that the macro and micro dynamics of complex systems are intimately linked, making for non-stationary non-ergodic behavior that cannot be reduced to…

[deleted]

Re: Extracting concepts from GPT-4

#73

Earlier quoted context omitted.

[flagged]

Is your argument that because AI can’t currently do the arbitrary things you wish it would do, it is therefore bullshit? This perspective discounts two important things: 1. All the things it can obviously do very well today 2. Future advancements to the tech (billions are pouring in, but this takes time to manifest in prod) I’m trying not to be one of the “guys” you’re talking about, but I just can’t comprehend your…

> 1. All the things it can obviously do very well today

I'm curious what those things are.

At least to me, it isn't obvious that LLMs solve any of their many applications from the past year "very well". I worry about failures (hallucinations, misinterpretation of prompts, regurgitation of incorrect facts, violation of copyright, and more). I don't have a good sense of when they fail, how often this happens, or how to identify these failures in scenarios where I'm not a domain expert.

But maybe some subset of these are solved problems, or at least problems that are actively being worked on for the next generation of models.

Yes, there are many other kinds of AI. Stockfish is better at chess than any human. But when you start talking about emergent behavior from machine learning, the failure modes are much harder to reason about.

Re: Extracting concepts from GPT-4

#74

Earlier quoted context omitted.

[flagged]

Is your argument that because AI can’t currently do the arbitrary things you wish it would do, it is therefore bullshit? This perspective discounts two important things: 1. All the things it can obviously do very well today 2. Future advancements to the tech (billions are pouring in, but this takes time to manifest in prod) I’m trying not to be one of the “guys” you’re talking about, but I just can’t comprehend your…

There are two mindsets at play here, the cynics vs the optimists. I’m an optimist to a fault by nature, but I also think there is a kind of Pascals bet to be played here.

If you bet sensibly on the current bubble/wave and turn out to be wrong - well you’re in the same place as everyone else with maybe some time and money lost.

But if you’re right things get interesting.

Re: Extracting concepts from GPT-4

#75
post #41

Earlier quoted context omitted.

LLMs aren't the only kind of AI, just one of the two current shiny kinds. If a "cure for cancer" (cancer is not just one disease so, unfortunately, that's not even as coherent a request as we'd all like it to be) is what you're hoping for, look instead at the stuff like AlphaFold etc.: https://en.wikipedia.org/wiki/AlphaFold I don't know how to tell where real science ends and PR bluster begins in such models, though…

[flagged]

> guys selling AI

Who?

Re: Extracting concepts from GPT-4

#76
post #47
post #12

Exciting to see this so soon after Anthropic's "Mapping the Mind of a Large Language Model" (under 3 weeks). I find these efforts really exciting; it is still common to hear people say "we have no idea how LLMs / Deep Learning works", but that is really a gross generalization as stuff like this shows. Wonder if this was a bit rushed out in response to Anthropic's release (as well as the departure of Jan Leike from Op…

> Wonder if this was a bit rushed out in response to Anthropic's release too lazy to dig up source but some twitter sleuth found that the first commit to the project was 6 months ago likely all these guys went to the same metaphorical SF bars, it was in the water

This project has been in the works for about a year. The initial commit to the public repo was not really closely related to this project, it was part of the release of the Transformer debugger, and the repo was just reused for this release.

Re: Extracting concepts from GPT-4

#77
post #12

Exciting to see this so soon after Anthropic's "Mapping the Mind of a Large Language Model" (under 3 weeks). I find these efforts really exciting; it is still common to hear people say "we have no idea how LLMs / Deep Learning works", but that is really a gross generalization as stuff like this shows. Wonder if this was a bit rushed out in response to Anthropic's release (as well as the departure of Jan Leike from Op…

We were planning to release the paper around this time independent of the other events you mention.

I think it is still predominantly accurate to say that we have no idea how LLMs work. SAEs might eventually change that, but there's still a long way to go.

Re: Extracting concepts from GPT-4

#78
post #10

Interesting, reminds me of similar work Anthropic did on Claude 3 Sonnet [0]. [0] https://transformer-circuits.pub/2024/scaling-monosemanticit...

The methods are the same, this is just OpenAI applying Anthropic's research to their own model.

The paper introduces substantial improvements over the methodology in the Anthropic SAE paper, and the research was done concurrently.

Re: Extracting concepts from GPT-4

#79
post #73

Earlier quoted context omitted.

Is your argument that because AI can’t currently do the arbitrary things you wish it would do, it is therefore bullshit? This perspective discounts two important things: 1. All the things it can obviously do very well today 2. Future advancements to the tech (billions are pouring in, but this takes time to manifest in prod) I’m trying not to be one of the “guys” you’re talking about, but I just can’t comprehend your…

> 1. All the things it can obviously do very well today I'm curious what those things are. At least to me, it isn't obvious that LLMs solve any of their many applications from the past year "very well". I worry about failures (hallucinations, misinterpretation of prompts, regurgitation of incorrect facts, violation of copyright, and more). I don't have a good sense of when they fail, how often this happens, or how to…

It's made me vastly more productive. It's a force multiplier in the areas I work, coding and content creation.

Re: Extracting concepts from GPT-4

#80
post #25

a.k.a, the same work as Anthropic, but with less interpretable and interesting features. I guess there won't be Golden Gate[0] GPT anytime soon. I mean, you just have to compare the couple of interesting features of the OpenAI feature browser [1] and the features of the Anthropic feature browser [2]. [0] https://twitter.com/AnthropicAI/status/1793741051867615494 [1] https://openaipublic.blob.core.windows.net/sparse-a…

Note that we focus on random positive activations, which are less susceptible to interpretability illusions than top activations (but also look less impressive as a result). We also provide access to random uncherrypicked features, whereas Anthropic does not. We made these choices deliberately to give as accurate an impression of autoencoder feature quality as possible.

Also note that GPT-4 is a more powerful model than Sonnet, which makes it harder to train autoencoders with the same quality features.

Post reply on HN