Live data from Hacker News

Continuous Diffusion Language Models (CDLM's)

sander.ai

11–20 of 55 posts

Re: Continuous Diffusion Language Models (CDLM's)

#11

> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022 I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. (The only exception I will make is encoder-decoder models…

> GPT2 was considered too dangerous to release

This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat.

Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies.

They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.

Re: Continuous Diffusion Language Models (CDLM's)

#12
post #11

> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022 I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. (The only exception I will make is encoder-decoder models…

> GPT2 was considered too dangerous to release This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat. Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies. They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.

My impression at the time was they were perhaps overly cautious but this was a bunch of researchers who wanted to self-regulate. Anthropic didn’t exist, deepmind was also much less product-focused and relatively cautious. Chinese models weren’t really a factor either.

Government regulation was really not in the picture either in 2020 or 2021 tbh. The government was still trying to beat a pandemic. Some of the Biden admin eventually wanted to but it wasn’t very serious.

Re: Continuous Diffusion Language Models (CDLM's)

#14
post #11

> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022 I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. (The only exception I will make is encoder-decoder models…

> GPT2 was considered too dangerous to release This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat. Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies. They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.

I agree that this should be something that researchers reflect on. GPT-2 is one of the primary models to research on nowadays, and many recent developments have come from studying it as a test bench.

Imagine if CRISPR was considered "too dangerous" to publish because of the potential ethical ramifications, and that only a special few should be aware. It is utter self-righteousness, and it is shameful behaviour. The world cannot adjust itself to what it cannot see, so you risk greater catastrophe by keeping it secret.

The open dissemination of knowledge at every increment is the only way for society to truly deal with what is to come.

Re: Continuous Diffusion Language Models (CDLM's)

#19
post #11

Earlier quoted context omitted.

> GPT2 was considered too dangerous to release This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat. Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies. They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.

My impression at the time was they were perhaps overly cautious but this was a bunch of researchers who wanted to self-regulate. Anthropic didn’t exist, deepmind was also much less product-focused and relatively cautious. Chinese models weren’t really a factor either. Government regulation was really not in the picture either in 2020 or 2021 tbh. The government was still trying to beat a pandemic. Some of the Biden a…

why are chinese models a consideration at all?

Re: Continuous Diffusion Language Models (CDLM's)

#20
I would love to see models that can think at different rates and also output a thinking scratchpad alongside output text instead of before all output.

Right now models need to rely on less legible compressed CoT to get high intelligence per token/step, but with diffusion they would just need to output more tokens per step instead.

Post reply on HN