Note that the distilled data is not even from the same "domain" of input data any more. They're basically adversarial inputs.
“Less than one”-shot learning
11–20 of 46 posts
Re: “Less than one”-shot learning
#12As an example, the media lab is still citing innovation with deep fakes, claiming entirely novel results people are shocked to see. They hype their own researchers even though there are kids on YouTube who that have been making similar content up to a year prior to Technology Review's publication.
I suspect they do the same with fields I'm less familiar with.
Re: “Less than one”-shot learning
#13Re: “Less than one”-shot learning
#14The title is misleading. The core technique still uses 60,000 images from MNIST, but 'distills' them into 10 images that contain the information from the original 60,000. The 10 'distilled' images look nothing like digits. Learning a complex model from 10 (later reduced to 2) 'distilled' number arrays is an interesting research idea, but it has little to do with reducing the size of the input dataset. Arguably the he…
It may mean we might finally have a method for reliably updating big neural networks instead of having to do continuous on the fly retraining. (Imagine future neural networks on your smart car having "upgrade packs" or country specific data that can be used to fine tune the main network in a matter of minutes) A high-level form of patch-and-diff for networks. There is probably a ML Ops startup opportunity somewhere i…
Re: “Less than one”-shot learning
#15Re: “Less than one”-shot learning
#16Earlier quoted context omitted.
mostly it seems that it tells you something sort of weird about how neural networks train. it’s not obvious that this should work, and that it can be made to work is interesting.
Agreed, the idea is interesting and worthy of further exploration. For starters, does it scales beyond MNIST or 'tiny synthetic datasets'?
Re: “Less than one”-shot learning
#17The title is misleading. The core technique still uses 60,000 images from MNIST, but 'distills' them into 10 images that contain the information from the original 60,000. The 10 'distilled' images look nothing like digits. Learning a complex model from 10 (later reduced to 2) 'distilled' number arrays is an interesting research idea, but it has little to do with reducing the size of the input dataset. Arguably the he…
Re: “Less than one”-shot learning
#18The title is misleading. The core technique still uses 60,000 images from MNIST, but 'distills' them into 10 images that contain the information from the original 60,000. The 10 'distilled' images look nothing like digits. Learning a complex model from 10 (later reduced to 2) 'distilled' number arrays is an interesting research idea, but it has little to do with reducing the size of the input dataset. Arguably the he…
The total entropy in 10 images (even carefully engineered ones) is very low in comparison to the full data-set.
Re: “Less than one”-shot learning
#19The title is misleading. The core technique still uses 60,000 images from MNIST, but 'distills' them into 10 images that contain the information from the original 60,000. The 10 'distilled' images look nothing like digits. Learning a complex model from 10 (later reduced to 2) 'distilled' number arrays is an interesting research idea, but it has little to do with reducing the size of the input dataset. Arguably the he…
Re: “Less than one”-shot learning
#20The title is misleading. The core technique still uses 60,000 images from MNIST, but 'distills' them into 10 images that contain the information from the original 60,000. The 10 'distilled' images look nothing like digits. Learning a complex model from 10 (later reduced to 2) 'distilled' number arrays is an interesting research idea, but it has little to do with reducing the size of the input dataset. Arguably the he…