Earlier quoted context omitted.
I think the point is that it's not a reconstruction. It's more like recognizing which letter of a thousand-letter alphabet is shown to the human after decoding their brain waves. Still impressive, but not really as impressive as visual reconstruction.
TBH, I was not impressed up until now, but given the videos I have in mind from people trying to use brain computing interfaces to type a text, now I'm impressed.
High-res image reconstruction with latent diffusion models from human brain
101–110 of 170 posts
Re: High-res image reconstruction with latent diffusion models from human brain
#102Are any of the example images novel, i.e. new to the model? Or is the model only reconstructing images it has already seen before? Either way, if I'm understanding right, it's very impressive. If the only input to the model (after training) is a fMRI reading, and from that it can reconstruct an image, at the very least that shows it can strongly correlate brain patterns back to the original image. It'd be even cooler…
There was a video many years ago (early 2010s?) demoing a similar technology, which would overlay and blend many images on top of each other to make a fuzzy image approximating what was actually being viewed. Edit: found it! https://youtu.be/nsjDnYxJ0bo
Re: High-res image reconstruction with latent diffusion models from human brain
#103Earlier quoted context omitted.
Nope. If you train a model where the input is an integer between 1 and 10, and the output is a specific image from a set of ten, the model will be able to get zero loss on the task. That is what's happening here.
Are you saying the demonstrated results are all in sample? Because this is definitely not true for out of sample data. And the GP comment implies that there is in fact a validation/holdout set.
Re: High-res image reconstruction with latent diffusion models from human brain
#104As people and groups increasingly move this direction do we think about vectors for abuse in 10, 20 or 50+ years? The human mind is considered the only place where we have true privacy. All these efforts are taking that away. At this rate all notions of privacy will soon be dead.
If this technology becomes accessible to courtrooms or police, they will use it. There will never be a way to encrypt thoughts.
Re: High-res image reconstruction with latent diffusion models from human brain
#105Earlier quoted context omitted.
People only think in words right before they say something, so I'm not sure how big a deal this is. I guess they'd be able to predict what I'm writing half a second before I write it? Would be useful if I lost the ability to write or speak, for whatever reason.
I'm "thinking in words" this entire thread as I read it. Do some people read without hearing the words in their head?
But if I'm just reading, I just take in chunks and phrases while constructing a meaning model. The sounds themselves (or even individual words) don't really enter into it.
Re: High-res image reconstruction with latent diffusion models from human brain
#106I immediately found the results suspect, and think I have found what is actually going on. The dataset it was trained on was 2770 images, minus 982 of those used for validation. I posit that the system did not actually read any pictures from the brains, but simply overfitted all the training images into the network itself. For example, if one looks at a picture of a teddy bear, you'd get an overfitted picture of anot…
Theoretically, the results would scale to more training images... we just need to fMRI all of LAION-5B. Easy peasy.
Re: High-res image reconstruction with latent diffusion models from human brain
#107Earlier quoted context omitted.
The only thing that's ethically questionable are humans themselves and such a device would likely do more to expose the unethical. Everybody remembers what happen when online dna hit mainstream.
I think only good thoughts so I'm not worried. Plus anyone with control of this technology is sure to have our best interests in mind.
Think good, happy thoughts. Happiness is mandatory. Being unhappy is treason. Treason is punishable by death.
Have a happy daycycle, citizen!
Re: High-res image reconstruction with latent diffusion models from human brain
#108Re: High-res image reconstruction with latent diffusion models from human brain
#109Earlier quoted context omitted.
I'm definitely not an expert in this subject, but even if the model is overfitted, doesn't the fact that it can pull out the similar images at all give credit to the idea that a larger, non-overfitted model could actually work as the paper describes? It means that there does exist some correlation between the shown subject, the captured fMRI data, and the resulting location in latent space.
The output part is basically nonsense. It would be more honest if the output was a text. E.g. "Teddybear" instead of a bad image of a random teddybear.
i.e) is there actually more information than a few bits encoding a crude object category, which stable diffusion then hallucinates the rest (/ uses to regurgitate an over-fit image)?
Or are there many bits, corresponding spatially to different regions of the stimulus - allowing for some meaningful degree of generalization.
Re: High-res image reconstruction with latent diffusion models from human brain
#110Earlier quoted context omitted.
I'm definitely not an expert in this subject, but even if the model is overfitted, doesn't the fact that it can pull out the similar images at all give credit to the idea that a larger, non-overfitted model could actually work as the paper describes? It means that there does exist some correlation between the shown subject, the captured fMRI data, and the resulting location in latent space.
Nope. If you train a model where the input is an integer between 1 and 10, and the output is a specific image from a set of ten, the model will be able to get zero loss on the task. That is what's happening here.
Although it seems they're only able to extract the subject of the brain activity, not any actual "pictures".