Fun times with energy-based models
mpmisko.github.io
Fun times with energy-based models
1–10 of 16 posts
Re: Fun times with energy-based models
#2Since we're on the subject, what are EBMs good for today?
Re: Fun times with energy-based models
#3I think the rationale for using tricks like score matching and contrastive divergence deserves a mention: the partition function is computationally expensive. Since we're on the subject, what are EBMs good for today?
- Simplicity and Stability: An EBM is the only object that needs to be trained and designed. Separate networks are not tuned to ensure balance.
- Sharing of Statistical Strength: Since the EBM is the only trained object, it requires fewer model parameters than approaches that use multiple networks.
- Adaptive Computation Time: Implicit sample generation is an iterative stochastic optimization process, which allows for a trade-off between generation quality and computation time.
- VAEs and flow-based models are bound by the manifold structure of the prior distribution and consequently have issues modelling discontinuous data manifolds, often assigning probability mass to areas unwarranted by the data. EBMs avoid this issue by directly modelling particular regions as high or lower energy.
- Compositionality: If we think of energy functions as costs for a certain goals or constraints, summation of two or more energies corresponds to satisfying all their goals or constraints.
Re: Fun times with energy-based models
#4I think the rationale for using tricks like score matching and contrastive divergence deserves a mention: the partition function is computationally expensive. Since we're on the subject, what are EBMs good for today?
Re: Fun times with energy-based models
#5I think the rationale for using tricks like score matching and contrastive divergence deserves a mention: the partition function is computationally expensive. Since we're on the subject, what are EBMs good for today?
Re: Fun times with energy-based models
#6I think the rationale for using tricks like score matching and contrastive divergence deserves a mention: the partition function is computationally expensive. Since we're on the subject, what are EBMs good for today?
p ∝ anchor_policy * exp(utility / temperature)
The utility is exactly the same as "energy". The article ignores entropy, but you can add in entropy regularization e.g. in soft actor-critic.
Re: Fun times with energy-based models
#7I think the rationale for using tricks like score matching and contrastive divergence deserves a mention: the partition function is computationally expensive. Since we're on the subject, what are EBMs good for today?
This paper lists the benefits in the introduction: https://proceedings.neurips.cc/paper_files/paper/2019/file/3... - Simplicity and Stability: An EBM is the only object that needs to be trained and designed. Separate networks are not tuned to ensure balance. - Sharing of Statistical Strength: Since the EBM is the only trained object, it requires fewer model parameters than approaches that use multiple networks. - Ada…
Re: Fun times with energy-based models
#8I think the rationale for using tricks like score matching and contrastive divergence deserves a mention: the partition function is computationally expensive. Since we're on the subject, what are EBMs good for today?