Live data from Hacker News

High-res image reconstruction with latent diffusion models from human brain

github.com

131–140 of 170 posts

Re: High-res image reconstruction with latent diffusion models from human brain

#131
post #28

Earlier quoted context omitted.

I'm definitely not an expert in this subject, but even if the model is overfitted, doesn't the fact that it can pull out the similar images at all give credit to the idea that a larger, non-overfitted model could actually work as the paper describes? It means that there does exist some correlation between the shown subject, the captured fMRI data, and the resulting location in latent space.

Nope. If you train a model where the input is an integer between 1 and 10, and the output is a specific image from a set of ten, the model will be able to get zero loss on the task. That is what's happening here.

But unless they tested this on a single human being; doesn't this mean that we can read brains (it's just this one particular reader is bad).

Re: High-res image reconstruction with latent diffusion models from human brain

#132

I immediately found the results suspect, and think I have found what is actually going on. The dataset it was trained on was 2770 images, minus 982 of those used for validation. I posit that the system did not actually read any pictures from the brains, but simply overfitted all the training images into the network itself. For example, if one looks at a picture of a teddy bear, you'd get an overfitted picture of anot…

I don’t get the criticism here. Normally I’d be the first to err on the side of skepticism, but this work seems above board.

I think the confusion is that this model is generating “teddy bear” internally, not a photo of a teddy bear. I.e. the diffusion part was added for flair, not to generate the details of the images that exist inside your mind. They could just as easily have run print(“teddy bear”), but they’re sending it to diffusion instead of printing it to console.

The fact that it can correctly discern between a dozen different outputs is pretty remarkable. And that’s all that this is showing. But that’s enough.

It’s not really a “gotcha” to say that it’s showing an image from the training set. They could have replaced diffusion with showing a static image of a teddy bear.

It sounds like this is many readers’ first time confronting the fact that scientists need to do these kinds of projects to get funding. As long as they’re not being intentionally deceptive, it seems fine. There’s a line between this and that ridiculous “rat brain flies plane” myth, and this seems above it.

Disclaimer: I should probably read the paper in detail before posting this, but the criticism of “the building looks like a training image” is mostly what I’m responding to. There are only so many topics one can think about, and having a machine draw a dog when I’m thinking about my dog Pip is some next-level sci-fi “we live in the future” stuff. Even if it doesn’t look like Pip, does it really matter?

Besides, it’s a matter of time till they correlate which parts of the brain are more prone to activating for specific details of the image you’re thinking about. Getting pose and color right would go a long way. So this is a resolution problem; we need more accurate brain sampling techniques, i.e. Neuralink. Then I’m sure diffusion will get a lot more of those details correct.

Re: High-res image reconstruction with latent diffusion models from human brain

#134
post #46

Earlier quoted context omitted.

The output part is basically nonsense. It would be more honest if the output was a text. E.g. "Teddybear" instead of a bad image of a random teddybear.

In this specific case I agree, since the model may be overfitted, it seems like it's currently just a glorified object classifier based on what was in the training data, but the fact that it works at all may indicate that the underlying idea has merit. They would probably have to train a much larger network to see if it's able to separate features distinctly enough using the input fMRI data to be useful.

My colleagues did the same, but with EEG. This makes the technique much more accessible: https://arxiv.org/abs/2302.10121

One open question in the field: how to assess the alignment of the AI outcomes across different methods?

Re: High-res image reconstruction with latent diffusion models from human brain

#135

Earlier quoted context omitted.

The problem is that it's impossible to know what is in the fMRI data and what is hallucinated by the reconstruction. In this case, the real bear has a blue ribbon and the "reconstructed" bear ha a red ribbon. Is the ribbon in the fMRI data and the computer choose the wrong color, or most of the images in the training set had ribbons and the computer just added one. Imagine this something like this is used in the futu…

> Imagine this something like this is used in the future to get something like https://en.wikipedia.org/wiki/Facial_composite . People may give too much importance to the details and arrest someone only because the computer imagined some detail, like the logo in the baseball cap. Wow, tech not working to tech might kill someone went super fast here.

In the real world when tech doesn't work people die.

OP is right to be concerned. This kind of tech (magickal mind-reading AI?!) is going to be bought up by security agencies, who wiil not understand its limitations and misuse it to accuse people of crimes they aren't related to.

There is ample precedent. Just for one recent example see plans to use an "AI lie detector" based on discredited pseudo-science at EU borders:

https://theintercept.com/2019/07/26/europe-border-control-ai...

Re: High-res image reconstruction with latent diffusion models from human brain

#137
post #99

Earlier quoted context omitted.

Why is it unethical to put chips in monkey brains?

Oooh I love moral philosophy. Here's one deontological perspective you could take: It is always wrong to cause unecessary suffering to others. Now, the subjective traits to be considered here are "unecessary" and "suffering." It used to be a common belief that animals lacked the capacity to suffer as humans do. They could feel pain, nocioception sure, but whether it caused complex psychological suffering (torment) us…

Nice analysis, shame about the cop-out ("I will leave the rest up to you!"). You're like- here's how to use morality, actually using it is left as an exercise to the reader :P

Re Jainism, adherents practice lacto-vegetarianism, but they, for example, don't eat tubers because they consider them too advanced, if I understand correctly. A deep respect for all forms of life is hard to get right in a world where every living thing eats some other living thing, or dies.

Re: High-res image reconstruction with latent diffusion models from human brain

#138
post #29

As people and groups increasingly move this direction do we think about vectors for abuse in 10, 20 or 50+ years? The human mind is considered the only place where we have true privacy. All these efforts are taking that away. At this rate all notions of privacy will soon be dead.

It feels like we’re at the point in the movie where someone travels back in time to warn humanity of the impending apocalypse soon to be unleashed by our insatiable appetite for technological advancement.

Re: High-res image reconstruction with latent diffusion models from human brain

#139

Earlier quoted context omitted.

Yeah, personally - and this might be a controversial opinion - I think most people who say they have an inner monologue are actually misattributing their experience of subvocalization while either reading text, or planning hypothetical conversations with other people (like a child might speak to their imaginary friend). It seems dubious to label that a monologue, because it depends on external stimuli (either the tex…

Wow, this is interesting, I thought pretty much everyone had an internal monologue. It feels like a colorblind person sharing the suspicion that everyone is just making up this other color spectrum.

Yeah, really interesting, who’s neurodivergent in this case?

I guess we’ll soon have a device that can find out!

Re: High-res image reconstruction with latent diffusion models from human brain

#140

Earlier quoted context omitted.

I'm "thinking in words" this entire thread as I read it. Do some people read without hearing the words in their head?

I don't have a word-based internal dialogue. When it's time to use words, like writing this post, sure, it's word time, so words are used. But if I'm just reading, I just take in chunks and phrases while constructing a meaning model. The sounds themselves (or even individual words) don't really enter into it.

That’s really fascinating and very different from my experience.

It’d be a really interesting project to measure and classify people’s individual thinking mechanisms. That daemon that seems to exist at the boundary of the conscious and unconscious.

Then again, maybe we wouldn’t want that as yet another data point to be bought and sold.

Post reply on HN