Live data from Hacker News

ChatGPT unexpectedly began speaking in a user's cloned voice during testing

arstechnica.com

141–150 of 164 posts

Re: ChatGPT unexpectedly began speaking in a user's cloned voice during testing

#141

Earlier quoted context omitted.

>there was simply text that is statistically likely to be arranged in that way This is wrong. An LLM can produce text that has never been arranged that way in training data.

Your statement doesn't contradict their statement. It produces text that is statistically likely to be arranged that way, but because we use non-deterministic sampling we get a wide variety of results that weren't necessarily in the training data.

No, I don't think they are talking about the decoding process there. If they were, it's a non statement. They are almost certainly talking about training data.

Re: ChatGPT unexpectedly began speaking in a user's cloned voice during testing

#142
post #78

Earlier quoted context omitted.

> LLM’s are good at detecting patterns and like to continue the pattern. They’re starting with autocomplete for voice and training it to do something else. This is a great summary of almost everything that goes wrong with LLM applications. LLMs are autocomplete machines, which is why GitHub Copilot is still the most reliably useful application of LLM tech out there. The further you get from autocomplete, the less rel…

You are making a very popular mistake: confusing the training procedure for LLMs (autocomplete) with "how they work"/their internal ontology (mostly unknown). When we teach children how to do arithmetic, we have them predict missing items in equations. We don't accuse them of "only doing autocomplete". The same applies for LLMs.

I am always confused when people compare LLMs with children - there is an obvious evolution of reasoning in children as the time passes by. This can be visible on a time span of months.

Using the same analogy: I dont see this in LLMs, I can speak with an LLM months and if we dont release a new version it is the same.

So they are not like children and I am not sure we really understand learning in children more so to the level that we can compare LLM with a child way of learning.

Re: ChatGPT unexpectedly began speaking in a user's cloned voice during testing

#143

Earlier quoted context omitted.

People fine tune LLMs for classification tasks. This is completely wrong.

You're just going around saying people are completely wrong without reading what they wrote or providing any justification for that claim. I'm not sure how to respond because your comment is a non sequitur.

The justification is that people are fine tuning LLMs for classification. They take out the last layer, replace it with a layer which maps to n classes instead of vocab_size, and the training data Y's aren't next word, they are a class label (I have a job which does binary classification, for example)

It's just completely wrong to say everything in LLM land is autocomplete. It's trivial and common to do the above.

Re: ChatGPT unexpectedly began speaking in a user's cloned voice during testing

#144

Earlier quoted context omitted.

Nope. If you actually tried those, you would quickly find out they don't work. It's actually really hard to clone a voice from a few seconds sample.

A few seconds, yeah. I've seen fairly convincing reproductions from 30 seconds of reading text though.

[deleted]

Re: ChatGPT unexpectedly began speaking in a user's cloned voice during testing

#145
post #80

Earlier quoted context omitted.

I'd respectfully disagree with this characterization of LLMs. While they certainly excel at pattern recognition, calling them mere "autocomplete machines" vastly undersells their capabilities. LLMs demonstrate complex reasoning, multi-modal understanding, and emergent behaviors that go well beyond simple pattern continuation. They've succeeded in areas like mathematical problem-solving, creative tasks, and various re…

I don't understand how anybody can still claim LLMs show "complex reasoning". It's been shown time and time again that they'll produce a correct chain of reasoning when given a problem (e.g. wolf, goat, cabbage crossing a river; 3 guards and a door; etc.) that is roughly similar to what's in the training data but will fail when given a sufficiently novel modification _while still producing output that is confidently…

GPT-3.5 is not the same thing. When people talk about LLMs having capabilities like complex reasoning, they're talking about current-gen models (eg Claude 3.5 Sonnet, GPT-4, Llama-405B).

Re: ChatGPT unexpectedly began speaking in a user's cloned voice during testing

#146
post #139

Earlier quoted context omitted.

It's feasible with advanced enough tech. The hard part isn't getting cameras to sign the files they produce. The hard part is to preserve the chain of custody as images are cropped, rescaled, recompressed etc. You can do it with tech like Intel SGX. But you also need serious defense of the camera platforms against hacking, of the CPUs, of the software stacks. And there's no demand. News orgs feel they should be impli…

They might use complicated tech because things are changing due to AI in a way that could degrade their brands trust. By having camera-signed videos, when folks create eg deep fakes of their news anchors/brands, there's a way for consumers to verify what's real. It lets their brands become more trustworthy. Yeah, preserving the chain of custody is hard. I was thinking there are a few options: (1) the signature of the…

I don't think there's any need for laws here. After all the whole point of a digital signature is that if it's removed or tampered with, that's detectable (assuming you expect it to be signed in the first place).

I've thought about this sort of approach many times in the past, and also done a lot of work with SGX and similar tech that implements remote attestation (RA). So I know how to build this sort of system conceptually. RA lets you do provable computations where the CPU can sign data, and the "public key" contains a hash of the program and data along with certificates that let you check the CPU was authentic. And it runs the program in a special CPU mode where the kernel and other things can't access the memory space. That's all you need to do verifiable computation.

So to preserve chain of custody you just have a set of filters or transforms that are SGX enclaves running ffmpeg or whatever, and each one attaches the attestation data to the output video which includes a hash of the input video. Then you gather up a certificate+signature over the original raw video from the camera (the cert is evidence the camera is authentic and the key is protected by the camera - you can get this from iPhone cameras), then an org certificate showing it came from a certain company, and then the attestation evidence for each transform. A set of scripts lets people verify all the evidence.

The problem is, after doing some business case analysis, I concluded it would only really be useful in some very small and specific cases:

1. Citizen journalists who are posting videos online that then get verified by news orgs. So the other way around. In this case the origin camera would use iPhone App Attestation to generate the source certificate, and all the fancy attested transform stuff isn't really important because it's the news org doing the transforms and doing the verifying.

2. Phone cam shots for insurance and other similar use cases. There is some business potential here, but it'd be sales force heavy as nobody knows the tech exists and deepfake fraud may not be a big enough problem for them to care (yet ...). If someone is looking for a startup idea, have this one for free.

3. Very new news companies that don't have any reputation yet and want to stand out from the crowd.

The thing is, for (3) or any place where a news org wants to increase the trust of the viewers, you don't need cryptography. That's just over-complicating things. You can just put a short random four letter code into the chyron that's unique to that particular shot you see on screen. Then on your website you have a page where the original unedited files can be downloaded by supplying the code. If you use cameras that produce cryptographic evidence like timestamps that's gravy, and for browny points you could publish video hashes into an unforgeable replicated log to stop you backdating footage. For most people that will be more than good enough. The sort of thing that causes people to lose trust in media is stuff not CNN broadcasting outright deepfakes, although that will happen eventually, but when they engage in selective editing, drop stories entirely, use archive footage and misrepresent it as something new etc.

The worst kind of fakery I've seen mainstream media engage in was Channel 4 UK's recent broadcast of a fake news segment, in which they "secretly filmed" someone who was pretending to be a racist Reform activist. People on X swiftly discovered that the person on-screen wasn't an activist at all but a professional actor, who had been putting on a fake accent the whole time (that he even advertised on his website). It looks for all the world like C4 broadcast entirely and truly fake news, knew they were doing it, and when they were called on it they just flat out refused to investigate knowing the British establishment was behind them all the way, as Reform is unpopular with the civil servant types who are supposed to police the media.

Unfortunately, for that kind of fakery there is no technological solution. Or, well, there is, but it's called social media + face recognition, and we already have it.

Re: ChatGPT unexpectedly began speaking in a user's cloned voice during testing

#147

Earlier quoted context omitted.

> The idea that there are sources you can trust to do your thinking for you is the more dangerous illusion in my opinion, and I’m not convinced that society will be harmed by poking some holes through it There is no alternative to this idea. It is completely impossible for an individual to possess all of the knowledge of everything that affects their lives. The only option for getting some of this information is goin…

> There is no alternative to this idea. I said do your thinking for you, not do your information gathering for you. I would suggest that you do not trust any single source to only ever tell you things that are true. If there’s a topic you want to know something about it’s a much better course of action to look at multiple different sources, and do your own thinking to come to your own conclusions. There are no author…

Oh, then I misunderstood. I thought you were against the notion of trusting a source for information, at all.

I very much agree with your actual point - that no source should be trusted absolutely, and that the only way to get a decent-to-solid idea on a topic is to consume multiple sources on that topic.

However, the problem is that even then people have relatively little time. It's important to have sources that one can rely on to be relatively accurate with a high probability, to get some vague idea about a topic you're not deeply invested in, but do care about somewhat. And I think this is where LLMs can hurt the most.

Re: ChatGPT unexpectedly began speaking in a user's cloned voice during testing

#148

Earlier quoted context omitted.

You're just going around saying people are completely wrong without reading what they wrote or providing any justification for that claim. I'm not sure how to respond because your comment is a non sequitur.

The justification is that people are fine tuning LLMs for classification. They take out the last layer, replace it with a layer which maps to n classes instead of vocab_size, and the training data Y's aren't next word, they are a class label (I have a job which does binary classification, for example) It's just completely wrong to say everything in LLM land is autocomplete. It's trivial and common to do the above.

That's still autocomplete. You use it by feeding in a context (all the text so far) and asking it to produce the next word (your fine-tuned classification token). The only difference is you don't ask for more tokens once you have one.

That's a very clever way of reducing a problem to autocomplete, but it doesn't change the paradigm.

Re: ChatGPT unexpectedly began speaking in a user's cloned voice during testing

#150
post #104
post #53

Earlier quoted context omitted.

> how well could they predict everything they are going to say? Maybe they can predict preferences, not everything we are going to say

Would you like to volunteer and test that theory? It’s stronger than a Turing test. You’d be competing against an AI to prove it’s really you and not an AI trained on what you have been posting. As determined by people who have been seeing you post this whole time. Or even more interestingly — after you hit send, we’ll compare to the 5 versions of what the AI predicted you’d say, and calculate the “loss”. If you lose…

Sure, I am doing some work, collecting my writings, preparing them for LLM. For example I could select my own text as positive and a generic one as negative for RLHF. The resulting model should predict when I'd like a piece of text. But it would need retraining to track my changes over time.
Post reply on HN