Live data from Hacker News

DALL-E Paper and Code

github.com

1–10 of 14 posts

Re: DALL-E Paper and Code

#4
Note that this is just the VAE component as used to help training and generating images, it will not let you create crazy images with natural language as used in the blog post (https://openai.com/blog/dall-e/).

More specifically from that link:

> [...] the image is represented using 1024 tokens with a vocabulary size of 8192.

> The images are preprocessed to 256x256 resolution during training. Similar to VQVAE, each image is compressed to a 32x32 grid of discrete latent codes using a discrete VAE1 that we pretrained using a continuous relaxation.

OpenAI also provides the encoder and decoder models and their weights.

However, with the decoder model, it's now possible to say train a text-encoding model to link up to that decoder (training on say an annotated image dataset) to get something close to the DALL-E demo OpenAI posted. Or something even better!

Re: DALL-E Paper and Code

#5
post #3

Has anyone tried this out?

A thread of examples from the provided notebook: https://twitter.com/ak92501/status/1364666124919447558

Note that these just demonstrate that arbitrary encoded input images match the decoded images, which is what would be expected from a VAE.

Re: DALL-E Paper and Code

#6

Note that this is just the VAE component as used to help training and generating images, it will not let you create crazy images with natural language as used in the blog post ( https://openai.com/blog/dall-e/ ). More specifically from that link: > [...] the image is represented using 1024 tokens with a vocabulary size of 8192. > The images are preprocessed to 256x256 resolution during training. Similar to VQVAE, eac…

Yeah unfortunately OpenAI has only released the weaker resnets and vision transformers they trained.

Some brilliant folks (Ryan Murdock [@advadnoun], Phil Wang [@lucidrains]) have tried to replicate their results with projects like big-sleep [0] with decent results, but even with this improved VAE we're still a ways from DALL-E quality results.

If anyone would like to play with the model check out either the Google Colab [1] (if you wanna run it on Google's cloud) or my site [2] (if you want a simplified UI).

[0]: https://github.com/lucidrains/big-sleep/

[1]: https://colab.research.google.com/drive/1MEWKbm-driRNF8PrU7o...

[2]: https://dank.xyz

Re: DALL-E Paper and Code

#7
post #3

Has anyone tried this out?

The repository linked is just a part of the entire model so it can't be used as is.

That said there is a completely implementation made by lucidrains[1] with some results, the only missing component now is the dataset.

[1]: https://github.com/lucidrains/DALLE-pytorch

Re: DALL-E Paper and Code

#10

Can someone explain what this even is for folks reading the description and going "this means nothing to me"?

About a month ago, OpenAI released information on their latest project -- a neural network which aims to generate images from text[1]. The results were impressive and the work received a lot of attention in the ML community. The repo linked in this post includes a small portion of the code used in the model. Perhaps I'm missing some context as well, but the code itself appears to be remarkably generic, not particularly interesting, and basically useless on its own. Perhaps the repo is in its early stages and more interesting developments may come later. I'd assume it's simply being upvoted because of all the OpenAI fanboys. Feel free to correct me if I'm mistaken and there is something remotely useful in this repo.

[1] https://openai.com/blog/dall-e/

Post reply on HN