Live data from Hacker News

Continuous Diffusion Language Models (CDLM's)

sander.ai

31–40 of 55 posts

Re: Continuous Diffusion Language Models (CDLM's)

#31

Earlier quoted context omitted.

why are chinese models a consideration at all?

Parent post is saying there was no community of peer models, just some people figuring it out as they went and with lots of slack to go slow if they wanted

Oh it's the typical china race thing, we can't let the antithesis of the us be better than the us etc. American exceptionalism

Re: Continuous Diffusion Language Models (CDLM's)

#32

Earlier quoted context omitted.

Parent post is saying there was no community of peer models, just some people figuring it out as they went and with lots of slack to go slow if they wanted

Oh it's the typical china race thing, we can't let the antithesis of the us be better than the us etc. American exceptionalism

The 'china' part is irrelevant. You're addicted to being mad. Parent post is just saying there were no other players

Re: Continuous Diffusion Language Models (CDLM's)

#33
post #11

Earlier quoted context omitted.

> GPT2 was considered too dangerous to release This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat. Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies. They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.

What are your thoughts on the HuggingFace incident?

Optimal next token generation disguised as something more.

Re: Continuous Diffusion Language Models (CDLM's)

#34

Earlier quoted context omitted.

Oh it's the typical china race thing, we can't let the antithesis of the us be better than the us etc. American exceptionalism

The 'china' part is irrelevant. You're addicted to being mad. Parent post is just saying there were no other players

There were other players, china included, but it wasn't relevant until LLMs became profitable? China is being singled out as usual because it's the antithesis of the us

Re: Continuous Diffusion Language Models (CDLM's)

#36

> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022 I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. (The only exception I will make is encoder-decoder models…

> I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked.

He's not talking about decoders, he's talking about auto-regression. Before ChatGPT, the dominant paradigm was fine-tuning BERT-like models.

> Before ChatGPT there really wasn’t much of a concept of pre-training and post-training.

Again, people spend years just post-training BERTs in various ways.

Re: Continuous Diffusion Language Models (CDLM's)

#37

Earlier quoted context omitted.

The 'china' part is irrelevant. You're addicted to being mad. Parent post is just saying there were no other players

There were other players, china included, but it wasn't relevant until LLMs became profitable? China is being singled out as usual because it's the antithesis of the us

But... there weren't. The other "players," China included, had duck all to show for it. Did they have the same potential capabilities? Sure, maybe. It wasn't just a profitability thing though. Their models sucked. China is being singled out in that comment because their models are exceptional, with very few competitors, so when you're talking about a history of the competition not being any good you _have_ to acknowledge one of the only good present-day models.

Re: Continuous Diffusion Language Models (CDLM's)

#39
post #36

> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022 I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. (The only exception I will make is encoder-decoder models…

> I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. He's not talking about decoders, he's talking about auto-regression. Before ChatGPT, the dominant paradigm was fine-tuning BERT-like models. > Before ChatGPT there really wasn’t much of a concept of pre-training and post-training. A…

GPT2 and 3 work via autoregression. In chat bots, decoders work via autoregression.

> people spend years just post-training BERTs in various ways

Yes, I was one of them. That’s not called “post-training” it’s called fine-tuning.

Re: Continuous Diffusion Language Models (CDLM's)

#40
post #11

> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022 I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked. (The only exception I will make is encoder-decoder models…

> GPT2 was considered too dangerous to release This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat. Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies. They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.

That's what's actually dangerous.

It's a different sort of danger. Knives, guns, etc are locally dangerous. The threat from AI is globally dangerous. Despite being a much lower risk of it happening, the blast radius is massively bigger, which means the overall danger is greater.

Also, maybe both things some be considered. Many countries manage the sale of knives, guns, etc. You don't need to accept that risk just because AI exists, or vice versa.

Post reply on HN