Live data from Hacker News

Attention Is Bayesian Inference

medium.com

31–33 of 33 posts

Re: Attention Is Bayesian Inference

#31
I've read through most of the first paper mentioned.

Here, the authors have taken set up two synthetic experiments where transformers have to learn the probability of observing events from a sampled from a "ground truth" Bayesian model. If the probability assigned by the transformers to the event space matches the Bayesian posterior predictive distribution, then the authors infer that the model is performing Bayesian inference for these tasks. Furthermore, they use this to argue that transformers are performing Bayesian inference in general (belief-propagation throughout layers).

The transformers are trained on thousands of different "ground truth" Bayesian models, each randomly initialized which means that there's no underlying signal to be learned besides the belief propagation mechanism itself. This makes me wonder if any sufficiently powerful maximum likelihood-based model would meet this criteria of "doing Bayesian inference" in this scenario.

The transformers in this paper do not intrinsically know to perform inference due to the fact that they're transformers. They perform inference because the optimal solution to the problems in the experiments is specifically to do inference, and transformers are powerful enough to model belief propagation. I find it hard to extrapolate that this is what is happening for LLMs, for example.

Re: Attention Is Bayesian Inference

#32
post #28

Earlier quoted context omitted.

That's just not true, and even if LLMs did introduce more errors than humans, if you can't trust the author to proof read a summary article about his own papers, then you shouldn't trust the papers either.

I agree with the latter. The fact that they use an LLM for the summary post without rewriting it in their own words already makes me not trust their papers.

Great, and I think that's incorrect, and only getting more incorrect every year, and perhaps you should consider trusting researchers in this field to know how and when to use their own tools correctly. I suppose that's all there is to say about that.

Re: Attention Is Bayesian Inference

#33
post #24
post #22

Just skimming, noticed lots of em dashes, interesting :).

It's so disappointing that this has become a meme. Lot's of people write with em-dashes. If you want to criticize the _writing_, then do so.

Yep. Love em dashes. It's a stupid tell.

A better tell IMO is an unnatural huge amount of editorialized h2s / h3s. Often they are overly lofty.

Post reply on HN