Embarrassingly simple self-distillation improves code generation
1–10 of 227 posts
Re: Embarrassingly simple self-distillation improves code generation
#2Sorry apple, SSD is already taken, you can't use that acronym.
Re: Embarrassingly simple self-distillation improves code generation
#3Re: Embarrassingly simple self-distillation improves code generation
#4> simple self-distillation (SSD): Sorry apple, SSD is already taken, you can't use that acronym.
Re: Embarrassingly simple self-distillation improves code generation
#5Re: Embarrassingly simple self-distillation improves code generation
#6I suppose we just don't have a deeper underlying theory to lean on and help us 'design' anything.
Re: Embarrassingly simple self-distillation improves code generation
#7[flagged]
Re: Embarrassingly simple self-distillation improves code generation
#8> simple self-distillation (SSD): Sorry apple, SSD is already taken, you can't use that acronym.
Consistency Preservation Update (CPU)
Guided Probability Update (GPU)
History-aware Distillation Driving (HDD)
Probability Smoothing Update (PSU)
Re: Embarrassingly simple self-distillation improves code generation
#9We really need to develop better tools to understand what's happening inside these NNs. Working with high-D spaces is not something we're good at, and we're basically throwing stuff at it and seeing if it sticks.
Re: Embarrassingly simple self-distillation improves code generation
#10Title should be: Simple Self-Distillation Improves Code Generation