Live data from Hacker News

Extracting concepts from GPT-4

openai.com

141–150 of 155 posts

Re: Extracting concepts from GPT-4

#141
post #136
post #115

Earlier quoted context omitted.

Which is so broad as to be unhelpful. We also know that petroleum mixed with air may be combusted to release energy; we needed to characterise this much better in order for the motor car to be distinguishable from a fuel-air bomb.

And that's exactly my point. Regulating the underlying tech is utterly pointless in this case - it's utterly harmless by itself.

That's exactly wrong, we know some things that can be expressed as "a sequence of tokens" are harmful and indeed have already made them crimes.

What we need is to characterise what is possible so we can skip the AI equivalent of Union Carbide in Bohopal.

Re: Extracting concepts from GPT-4

#142

This is super cool, it feels like going in the direction of the "deep"/high level type of semantic searching I've been waiting for. I like their examples of basically filtering documents for the "concept" of price increases, or even something as high level as a rhetorical question I wonder how this compares to training/fine tuning a model on examples of rhetorical questions and asking it to find it in a given documen…

Exa is trying to do this. I've found some sort of interesting stuff this way but it honestly doesn't feel quite good enough yet to me. https://exa.ai/search?c=all

Exa's top hit for this article:

https://openai.com/index/language-models-can-explain-neurons...

Re: Extracting concepts from GPT-4

#143
post #12

Exciting to see this so soon after Anthropic's "Mapping the Mind of a Large Language Model" (under 3 weeks). I find these efforts really exciting; it is still common to hear people say "we have no idea how LLMs / Deep Learning works", but that is really a gross generalization as stuff like this shows. Wonder if this was a bit rushed out in response to Anthropic's release (as well as the departure of Jan Leike from Op…

> Mapping the Mind of a Large Language Model The fact that a paper is implying a LLM has a mind doesn't exactly bode well for the people who wrote it, not to mention the continued meaningless babbling about "safety". It'd also be nice if they could show their work so we could replicate it. Still, not shabby for an ad!

Well - what is a mind exactly? We don't really have a good definition for a human mind. Not sure we should be claiming domain over the term. It's not a terrible shorthand for discussing something that reads and responds as if it had some kind of mind - whether technically true or not (which we honestly don't know).

Re: Extracting concepts from GPT-4

#144

Earlier quoted context omitted.

> Mapping the Mind of a Large Language Model The fact that a paper is implying a LLM has a mind doesn't exactly bode well for the people who wrote it, not to mention the continued meaningless babbling about "safety". It'd also be nice if they could show their work so we could replicate it. Still, not shabby for an ad!

Well - what is a mind exactly? We don't really have a good definition for a human mind. Not sure we should be claiming domain over the term. It's not a terrible shorthand for discussing something that reads and responds as if it had some kind of mind - whether technically true or not (which we honestly don't know).

> It's not a terrible shorthand for discussing something that reads and responds as if it had some kind of mind

I really don't see it like that—it has very little memory, it has no ability to introspect before "choosing" what to say, no awareness of the concept of the coherency of statements (i.e. whether or not it's saying things that directly contradict its training), seems to have little sense of non-pattern-driven computation beyond what token patterns can encode at a surface level (e.g. of course it knows 1 + 1 = 2, but does it recognize odd notation/can it recognize and analyze arbitrary statements? of course not). I fully grant it is compelling evidence we can replicate many brain-like processes with software neural nets, but that's an entirely different thing than raising it to a level of thought or consciousness or self-awareness (which I argue is necessary in order to appropriately issue coherent statements, as perspective is a necessary thing to address even when attempting to make factual statements), but it strikes me as a lot closer to an analogy for a potential constituent component of a mind rather than a mind per se.

Re: Extracting concepts from GPT-4

#145
post #141
post #136

Earlier quoted context omitted.

And that's exactly my point. Regulating the underlying tech is utterly pointless in this case - it's utterly harmless by itself.

That's exactly wrong, we know some things that can be expressed as "a sequence of tokens" are harmful and indeed have already made them crimes. What we need is to characterise what is possible so we can skip the AI equivalent of Union Carbide in Bohopal.

Yes, but we also know that a knife can be used to slice vegetables or stab people, and we still allow knives. I can go to Google right now and easily find out how to make Sarin or ricin at home. Are you suggesting that we should ban Google Search because of that?

Re: Extracting concepts from GPT-4

#146
post #145
post #141

Earlier quoted context omitted.

That's exactly wrong, we know some things that can be expressed as "a sequence of tokens" are harmful and indeed have already made them crimes. What we need is to characterise what is possible so we can skip the AI equivalent of Union Carbide in Bohopal.

Yes, but we also know that a knife can be used to slice vegetables or stab people, and we still allow knives. I can go to Google right now and easily find out how to make Sarin or ricin at home. Are you suggesting that we should ban Google Search because of that?

> Yes, but we also know that a knife can be used to slice vegetables or stab people, and we still allow knives.

I'm from the UK originally, and guess what.

Also missing the point, given stabbing is a crime; what's the AI equivalent of a stabbing? Does anyone on the planet know?

> I can go to Google right now and easily find out how to make Sarin or ricin at home. Are you suggesting that we should ban Google Search because of that?

Google search has restrictions on what you can search for, and on what results it can return. The question is where to set those thresholds, those limits — and politicians do regularly argue about this for all kinds of reasons much weaker than actual toxins. The current fight in the US over Section 230 looks like it's about what can and can't be done and by whom and who is considered liable for unlawful content, despite the USA being (IMO) the global outlier in favour of free speech due to its maximalist attitude and constitution.

People joke about getting on watchlists due to their searches, and at least one YouTuber I follow has had agents show up to investigate their purchases.

Facebook got flack from the UN because they failed to have appropriate limits on their systems, leading to their platform being used to orchestrate the (still ongoing) genocide in Myanmar.

What's being asked for here is not the equivalent of "ban google search", it's "figure out the extent to which we need an equivalent of Section 230, an equivalent of law enforcement cooperation, an equivalent of spam filtering, an equivalent of the right to be forgotten, of etc." — we don't even have the questions yet, we have the analogies, that's all, and analogies aren't good enough regardless of if the system that you fear might do wrong is an AI or a regulatory body.

Re: Extracting concepts from GPT-4

#147

Earlier quoted context omitted.

When you can ask an AI for an entire book with no errors in the output… god that would be a huge token model

Copyright violation isn't just when you can output 100% exact copies of books. And don't forget, they also violated copyright internally billions of times during training. If any of us had been caught making copies of corporate-owned content for AI training use five years ago, we'd be in for zillion-dollar lawsuits that would make any grandma who downloaded a song from Napster blush.

If you copy your cd for backup with no resale future, no one would waste time to sue you.

Re: Extracting concepts from GPT-4

#148
post #146
post #145

Earlier quoted context omitted.

Yes, but we also know that a knife can be used to slice vegetables or stab people, and we still allow knives. I can go to Google right now and easily find out how to make Sarin or ricin at home. Are you suggesting that we should ban Google Search because of that?

> Yes, but we also know that a knife can be used to slice vegetables or stab people, and we still allow knives. I'm from the UK originally, and guess what. Also missing the point, given stabbing is a crime; what's the AI equivalent of a stabbing? Does anyone on the planet know? > I can go to Google right now and easily find out how to make Sarin or ricin at home. Are you suggesting that we should ban Google Search be…

What, you aren’t allowed to own kitchen knives? Or Google search somehow doesn’t return the chemical processes to make Sarin? Come on now.

Re: Extracting concepts from GPT-4

#149
post #148
post #146

Earlier quoted context omitted.

> Yes, but we also know that a knife can be used to slice vegetables or stab people, and we still allow knives. I'm from the UK originally, and guess what. Also missing the point, given stabbing is a crime; what's the AI equivalent of a stabbing? Does anyone on the planet know? > I can go to Google right now and easily find out how to make Sarin or ricin at home. Are you suggesting that we should ban Google Search be…

What, you aren’t allowed to own kitchen knives? Or Google search somehow doesn’t return the chemical processes to make Sarin? Come on now.

> What, you aren’t allowed to own kitchen knives?

You're not allowed to be in possession of a knife in public without a good reason.

https://www.gov.uk/buying-carrying-knives

You may think the UK government is nuts (I do, I left due to an unrelated law), but it is what it is.

> Or Google search somehow doesn’t return the chemical processes to make Sarin?

You're still missing the point of everything I've said if you think that's even a good rhetorical question.

I have no idea if that's me giving bad descriptions, or you being primed with the exact false world model I'm trying to convince you to change from.

Hill climbing sometimes involves going down one local peak before you can climb the global.

Again, and I don't know how to make this clearer, I am not calling for an undifferentiated ban on all AI just because they can be used for bad ends, I'm saying that we need to figure out how to even tell which uses are even the bad ones.

Your original text was:

> We know exactly what the system is capable of doing. It’s capable of outputting tokens which can then be converted into text

Well, we know exactly what a knife is capable of doing.

Does that knowledge mean we allow stabbing? Of course not!

What's the AI equivalent of a stabbing? Nobody knows.

Re: Extracting concepts from GPT-4

#150

Earlier quoted context omitted.

Copyright violation isn't just when you can output 100% exact copies of books. And don't forget, they also violated copyright internally billions of times during training. If any of us had been caught making copies of corporate-owned content for AI training use five years ago, we'd be in for zillion-dollar lawsuits that would make any grandma who downloaded a song from Napster blush.

If you copy your cd for backup with no resale future, no one would waste time to sue you.

Because they wouldn't catch me. But if they did, especially if they caught me making a copy of every CD at the CD store as a backup, especially if they caught me making a copy of every bootleg CD I could get my hands on (as a backup), I'd be in big trouble.

Did you know a lot of LLM training data is scraped from illegal pirate libraries such as Anna's Archive?

Post reply on HN