Live data from Hacker News

Continuous Diffusion Language Models (CDLM's)

sander.ai

41–50 of 55 posts

Re: Continuous Diffusion Language Models (CDLM's)

#41
post #11

> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022 I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. (The only exception I will make is encoder-decoder models…

> GPT2 was considered too dangerous to release This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat. Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies. They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.

I gave up on gun control as a political question when I realized you can buy a flamethrower at Home Depot

Re: Continuous Diffusion Language Models (CDLM's)

#42
post #11

Earlier quoted context omitted.

> GPT2 was considered too dangerous to release This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat. Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies. They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.

I gave up on gun control as a political question when I realized you can buy a flamethrower at Home Depot

[deleted]

Re: Continuous Diffusion Language Models (CDLM's)

#43
I love the idea of diffusion language models. I think they are potentially superior to autoregressive models, in the way that practically, considering whole systems tends to beat optimizing any one individual detail. To me, the way autoregressive models are sampled feels very fundamentally limited, and diffusion feels much more coherent by comparison.

Re: Continuous Diffusion Language Models (CDLM's)

#44
post #38

Conspiracy time: The continuous diffusion model has fruit. Frontier labs were trying to bury it with discrete diffusion distraction, because it is what they're using internally. Momentum wins, in the end.

No conspiracy. Diffusion is just a more finicky, more expensive way of generating the same tokens as autoregressive decoding.

Re: Continuous Diffusion Language Models (CDLM's)

#45
post #36

Earlier quoted context omitted.

> I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. He's not talking about decoders, he's talking about auto-regression. Before ChatGPT, the dominant paradigm was fine-tuning BERT-like models. > Before ChatGPT there really wasn’t much of a concept of pre-training and post-training. A…

GPT2 and 3 work via autoregression. In chat bots, decoders work via autoregression. > people spend years just post-training BERTs in various ways Yes, I was one of them. That’s not called “post-training” it’s called fine-tuning.

Asking out of curiosity, because I have limited experience in that domain.

I thought that fine-tuning was changing the weights in the model, not the embedding? Or did I misunderstand?

Re: Continuous Diffusion Language Models (CDLM's)

#46
post #11

> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022 I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. (The only exception I will make is encoder-decoder models…

> GPT2 was considered too dangerous to release This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat. Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies. They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.

As I understood it at the time, the rationale was that GPT2 made it too easy to produce content at scale and could easily be used for propaganda, phishing or other kinds of cons.

Which turned out to be true.

Re: Continuous Diffusion Language Models (CDLM's)

#47
post #11

> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022 I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. (The only exception I will make is encoder-decoder models…

> GPT2 was considered too dangerous to release This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat. Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies. They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.

> Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies.

You are hyperbolically judging a post-hoc cherry-picked subset of a great spectrum of concern, with hindsight experience nobody had at the time.

Safely navigating the future isn't an accuracy contest. Risk mitigation has to account for the distribution of costs for different signs of error, where novelty, uncertainty, and any potential for compounding effects all greatly multiply the need for hedging.

Re: Continuous Diffusion Language Models (CDLM's)

#48

> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022 I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. (The only exception I will make is encoder-decoder models…

I suppose what I was trying to say is that in the research community, people were still a lot more willing to entertain alternative modelling paradigms for language at that point, there was much less of a monoculture than there is today. It was really only after ChatGPT that non-autoregressive language modelling came to be seen as a fringe pursuit. Of course that is a highly subjective assessment, and my perspective is inevitably coloured by my own research interests at the time, and those of the people around me.

For what it's worth, I don't believe post-training diffusion models (of language or otherwise) is uniquely difficult, just a lot less thoroughly explored so far.

Re: Continuous Diffusion Language Models (CDLM's)

#50
post #11

> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022 I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. (The only exception I will make is encoder-decoder models…

> GPT2 was considered too dangerous to release This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat. Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies. They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.

They were worried it would make it easy to generate a flood of misinformation. They were correct.
Post reply on HN