Live data from Hacker News

Extracting concepts from GPT-4

openai.com

101–110 of 155 posts

Re: Extracting concepts from GPT-4

#101
post #12

Exciting to see this so soon after Anthropic's "Mapping the Mind of a Large Language Model" (under 3 weeks). I find these efforts really exciting; it is still common to hear people say "we have no idea how LLMs / Deep Learning works", but that is really a gross generalization as stuff like this shows. Wonder if this was a bit rushed out in response to Anthropic's release (as well as the departure of Jan Leike from Op…

From the article: "We currently don't understand how to make sense of the neural activity within language models." "Unlike with most human creations, we don’t really understand the inner workings of neural networks." "The [..] networks are not well understood and cannot be easily decomposed into identifiable parts" "[..] the neural activations inside a language model activate with unpredictable patterns, seemingly re…

Could there also be a “legal hedging” reason for why you would release a paper like this?

By reaffirming that “we don’t know how this works, nobody does” it’s easier to avoid being charged with copyright infringement from various actors/data sources that have sued them.

Re: Extracting concepts from GPT-4

#102

Earlier quoted context omitted.

From the article: "We currently don't understand how to make sense of the neural activity within language models." "Unlike with most human creations, we don’t really understand the inner workings of neural networks." "The [..] networks are not well understood and cannot be easily decomposed into identifiable parts" "[..] the neural activations inside a language model activate with unpredictable patterns, seemingly re…

I read this as "we have not built up tools / math to understand neural networks as they are new and exciting" and not as "neural networks are magical and complex and not understandable because we are meddling with something we cannot control". A good example would be planes - it took a long while to develop mathematical models that could be used to model behavior. Meanwhile practical experimentation developed decent…

> we don't have math / models yet that can explain/model their behavior...

So, what you're saying is we don't know how they work yet? It's not that deep.

Re: Extracting concepts from GPT-4

#103
post #12

Exciting to see this so soon after Anthropic's "Mapping the Mind of a Large Language Model" (under 3 weeks). I find these efforts really exciting; it is still common to hear people say "we have no idea how LLMs / Deep Learning works", but that is really a gross generalization as stuff like this shows. Wonder if this was a bit rushed out in response to Anthropic's release (as well as the departure of Jan Leike from Op…

Indeed, and the very last section about how they’ve now “open sourced” this research is also a bit vague. They’ve shared their research methodology and findings… But isn’t that obligatory when writing a public paper?

https://github.com/openai/sparse_autoencoder

They actually open sourced it, for GPT-2 which is an open model.

Re: Extracting concepts from GPT-4

#104
post #91

Earlier quoted context omitted.

We know exactly what the system is capable of doing. It’s capable of outputting tokens which can then be converted into text.

And social media manipulation is just registers and bytes, wait no, sand and electrons.

Just because you can do something with technology doesn't mean the problem is technology itself. It's like newspapers. Printing them I technology and allows all kind of things. If you're of the authoritarian mindset, you'll want to control it all out of some stated fear, but you can do that for everything.

Re: Extracting concepts from GPT-4

#105
post #25

a.k.a, the same work as Anthropic, but with less interpretable and interesting features. I guess there won't be Golden Gate[0] GPT anytime soon. I mean, you just have to compare the couple of interesting features of the OpenAI feature browser [1] and the features of the Anthropic feature browser [2]. [0] https://twitter.com/AnthropicAI/status/1793741051867615494 [1] https://openaipublic.blob.core.windows.net/sparse-a…

yeah this one is much less presentable than Anthropic's work. It sure looks bad on them to be compared so poorly like this.

Re: Extracting concepts from GPT-4

#106
post #80
post #25

a.k.a, the same work as Anthropic, but with less interpretable and interesting features. I guess there won't be Golden Gate[0] GPT anytime soon. I mean, you just have to compare the couple of interesting features of the OpenAI feature browser [1] and the features of the Anthropic feature browser [2]. [0] https://twitter.com/AnthropicAI/status/1793741051867615494 [1] https://openaipublic.blob.core.windows.net/sparse-a…

Note that we focus on random positive activations, which are less susceptible to interpretability illusions than top activations (but also look less impressive as a result). We also provide access to random uncherrypicked features, whereas Anthropic does not. We made these choices deliberately to give as accurate an impression of autoencoder feature quality as possible. Also note that GPT-4 is a more powerful model t…

[deleted]

Re: Extracting concepts from GPT-4

#107
post #45
post #10

Interesting, reminds me of similar work Anthropic did on Claude 3 Sonnet [0]. [0] https://transformer-circuits.pub/2024/scaling-monosemanticit...

I feel the webpage strongly hints that sparse autoencoders were invented by OpenAI for this project. Very weird that they don't cite this in their webpage and instead bury the source in their paper.

Nahhh, that's the tried-and-true Apple approach to marketing, and OpenAI is well positioned to adopt it for themselves. They act like they invented transformers as much as Apple acts like they invented the smartphone.

Re: Extracting concepts from GPT-4

#108
post #10

Interesting, reminds me of similar work Anthropic did on Claude 3 Sonnet [0]. [0] https://transformer-circuits.pub/2024/scaling-monosemanticit...

The methods are the same, this is just OpenAI applying Anthropic's research to their own model.

I'm the research lead of Anthropic's interpretability team. I've seen some comments like this one, which I worry downplay the importance of @leogao et al's paper due to the similarity of ours. I think these comments are really undervaluing Gao et al's work.

It's not just that this is contemporaneous work (a project like this takes many months at the very least), but also that it introduces a number of novel contributions like TopK activations and new evaluations. It seems very possible that some of these innovations will be very important for this line of work going forward.

More generally, I think it's really unfortunate when we don't value contemporaneous work or replications. Prior to this paper, one could have imagined it being the case that sparse autoencoders worked on Claude due some idiosyncracy, but wouldn't work on other frontier models for some reason. This paper can give us increased confidence that they work broadly, and that in itself is something to celebrate. It gives us a more stable foundation to build on.

I'm personally really grateful to all the authors of this paper for their work pushing sparse autoencoders and mechanistic interpretability forward.

Re: Extracting concepts from GPT-4

#109
post #49

Earlier quoted context omitted.

Hehe, related to this, someone created a "book4" dataset and put it on torrent websites. I don't think it's being used in any major LLMs, but the future "piracy" community intersection with AI is going to be exciting. Watching the cyberpunk world that all of my favorite literature predicted slowly come to our world is fun indeed.

i think you mean @sillysaurus' books3? not books4?

https://news.ycombinator.com/item?id=40405443

Re: Extracting concepts from GPT-4

#110
post #100

Earlier quoted context omitted.

> Is your argument that because AI can’t currently do the arbitrary things you wish it would do, it is therefore bullshit? It is sold as a research tool, but it cannot be trusted to return facts, because it will happily recombine disconnected pieces of data. AI cannot tell truth from lies, it is good at constructing output that looks like an answer but it does not care about the factual correctness. Google search res…

Your first section is very much a limit of LLMs, but again, that's not all AI — if you want an AI to play chess, and you want to actually win, you use Stockfish or AlphaZero, because if you use an LLM it will perform illegal moves.

Why would I want to use an AI to win a game of chess? Where's the fun and challenge in it? "Go win me a chess tournament" is the wish nobody has unless we are talking about someone who wants to pretend to be a chess master. It's still a small market. Examples like these are very common in the AI community, they are solutions to problems nobody has.
Post reply on HN