Live data from Hacker News

Extracting concepts from GPT-4

openai.com

41–50 of 155 posts

Re: Extracting concepts from GPT-4

#41

Earlier quoted context omitted.

From the article: "We currently don't understand how to make sense of the neural activity within language models." "Unlike with most human creations, we don’t really understand the inner workings of neural networks." "The [..] networks are not well understood and cannot be easily decomposed into identifiable parts" "[..] the neural activations inside a language model activate with unpredictable patterns, seemingly re…

Not holding my breath for that hallucinated cure for cancer then.

LLMs aren't the only kind of AI, just one of the two current shiny kinds.

If a "cure for cancer" (cancer is not just one disease so, unfortunately, that's not even as coherent a request as we'd all like it to be) is what you're hoping for, look instead at the stuff like AlphaFold etc.: https://en.wikipedia.org/wiki/AlphaFold

I don't know how to tell where real science ends and PR bluster begins in such models, though I can say that the closest I've heard to a word against it is "sure, but we've got other things besides protein folding to solve", which is a good sign.

(I assume AlphaFold is also a mysterious black box, and that tools such as the one under discussion may help us demystify it too).

Re: Extracting concepts from GPT-4

#42
post #26
post #17

Earlier quoted context omitted.

Notice that most of the examples have none of the green highlight counter, which is shown for > small losses. KEEPING SCORE: The Dow Jones industrial average rose 32 points, or 0.2 percent, to 18,156 as of 3:15 p.m. Eastern time. The Standard & Poor’s ... OMAHA, Neb. (AP) — Warren Buffett’s company has bought nearly the other sentences are in contrast to show how specific this neuron is.

The highlights are better visible in this visualisation: https://openaipublic.blob.core.windows.net/sparse-autoencode... There are also many top activations not showing increases, e.g. > 0.06 of a cent to 90.01 cents US.↵↵U.S. indexes were mainly lower as the Dow Jones industrials lost 21.72 points to 16,329.53, the Nasdaq was up 11.71 points at 4,318.9 and the S&P 500 (Highlight on the first comma.)

[deleted]

Re: Extracting concepts from GPT-4

#43
post #27
post #24

Earlier quoted context omitted.

I saw that, but the language makes me think it's not quite the same as what's really being used? "how a piece of text might be tokenized by a language model" "It's important to note that the exact tokenization process varies between models."

That's why they have buttons to choose which model's tokenizer to use.

Yes, thank you, I understand that part.

It's the might condition in the description that makes me think the results might not be the exact same as what's used in the live models.

Re: Extracting concepts from GPT-4

#45
post #10

Interesting, reminds me of similar work Anthropic did on Claude 3 Sonnet [0]. [0] https://transformer-circuits.pub/2024/scaling-monosemanticit...

I feel the webpage strongly hints that sparse autoencoders were invented by OpenAI for this project.

Very weird that they don't cite this in their webpage and instead bury the source in their paper.

Re: Extracting concepts from GPT-4

#46
post #35

This is interesting: > Autoencoder family > Note: Only 65536 features available. Activations shown on The Pile (uncopyrighted) instead of our internal training dataset. So, the Pile is uncopyrighted, but the internal training dataset is copyrighted? Copyrighted by whom? Huh?

Hehe, related to this, someone created a "book4" dataset and put it on torrent websites. I don't think it's being used in any major LLMs, but the future "piracy" community intersection with AI is going to be exciting.

Watching the cyberpunk world that all of my favorite literature predicted slowly come to our world is fun indeed.

Re: Extracting concepts from GPT-4

#47
post #12

Exciting to see this so soon after Anthropic's "Mapping the Mind of a Large Language Model" (under 3 weeks). I find these efforts really exciting; it is still common to hear people say "we have no idea how LLMs / Deep Learning works", but that is really a gross generalization as stuff like this shows. Wonder if this was a bit rushed out in response to Anthropic's release (as well as the departure of Jan Leike from Op…

> Wonder if this was a bit rushed out in response to Anthropic's release

too lazy to dig up source but some twitter sleuth found that the first commit to the project was 6 months ago

likely all these guys went to the same metaphorical SF bars, it was in the water

Re: Extracting concepts from GPT-4

#48
post #31
post #10

Interesting, reminds me of similar work Anthropic did on Claude 3 Sonnet [0]. [0] https://transformer-circuits.pub/2024/scaling-monosemanticit...

Someone mentioned that this took almost as much compute to train as the original model.

source please!

Re: Extracting concepts from GPT-4

#49
post #35

This is interesting: > Autoencoder family > Note: Only 65536 features available. Activations shown on The Pile (uncopyrighted) instead of our internal training dataset. So, the Pile is uncopyrighted, but the internal training dataset is copyrighted? Copyrighted by whom? Huh?

Hehe, related to this, someone created a "book4" dataset and put it on torrent websites. I don't think it's being used in any major LLMs, but the future "piracy" community intersection with AI is going to be exciting. Watching the cyberpunk world that all of my favorite literature predicted slowly come to our world is fun indeed.

i think you mean @sillysaurus' books3? not books4?
Post reply on HN