Live data from Hacker News

Dall-E 2

openai.com

91–100 of 511 posts

Re: Dall-E 2

#91

Earlier quoted context omitted.

Oh, no, the society! A picture of Joe Biden killing a priest! Society didn't collapse after photoshop. "Responsibility to society" is such a catch-all excuse.

You missed half of my note. An artist can say "no". A machine cannot. If you lower the barrier and allow anything, then you are responsible for the outcome. OpenAI rightfully took a responsible angle.

there will and are million ways to create a photorealistic picture of Joe Biden killing a priest using modern tools, and absolutely nothing will happen if someone did.

We've been through this many times, with books, with movies, with video games, with Internet. If it *can* be used for porn / violence etc., it will be, but it won't be the main use case and it won't cause some societal upheaval. Kids aren't running around pulling cops out of cars GTA-style, Internet is not ALL PORN, there is deepfake porn, but nobody really cares, and so on. There are so many ways to feed those dark urges that censorship does nothing except prevent normal use cases that overlap with the words "violence" or "sex" or "politics" or whatever the boogeyman du jour is.

Re: Dall-E 2

#92
I'm curious, is this something feasible to train (and inference) on a consumer level machine, or this is something can only be done by institutes?

Re: Dall-E 2

#93
post #78

Earlier quoted context omitted.

This feels unnecessarily hostile. I've felt a similar tinge of disappointment upon reading that paragraph, despite the fact that I somehow knew it was "their service, their call" without you being there to spell it out for me. It's also incredibly shortsighted of you to assume that people are interested in exploring this tool only as a means of generating art that they cannot themselves do. Eg. I myself am a software…

Quoted post unavailable.

> this service should be provided to me and if it isn't done how I want it that's infringing on me somehow

That is an extremely uncharitable interpretation of:

> I wish I could have a version with the training wheels taken off.

Re: Dall-E 2

#94
post #10

A friend of mine was studying graphic design, but became disillusioned and decided to switch to frontend programming after he graduated. His thesis advisor said he should be cautious, because automation/AI will soon take the jobs of programmers, implying that graphic design is a safer bet in this regard. Looks like his advisor is a few years from being proven horribly wrong.

I think designers are becoming more valuable than ever. Designers can better help train the AI on what actually looks good, designers will (probably) always have a more intuitive understanding of UI/UX, designers can better implement the work the AI actually produces, and designers can coordinate designs across multiple different mediums and platforms.

Additionally, the rise of no-code development is just extending the functionality of designers. I didn't take design seriously (as a career choice) growing up because I didn't see a future in it, now it pays my bills and the demand for my services just grows by the day.

Similar argument to make with chess AI: it didn't make chess players obsolete, it made them stronger than ever.

Re: Dall-E 2

#95
post #19

Am I the only one to think that the AI world is divided into 2 groups: 1. Deepmind, who solved go, protein folding, and that seems really onto something. 2. Everyone else, spending billions to build machines that draw astronauts on unicorns, and smartish bot toys.

Your second group represents the core "inner loop" of about a thousand revolutionary applications. Take the basic capability of translating image->text->speech (and the reverse), install it on a wearable device that can "see" an environment, and add domain-specific agents. From this setup, you're not too far away from having an AI that can whisper guidance into your ear like a co-pilot, enabling scenarios like:

1. step-by-step guidance for a blind person navigating the use of a public restroom.

2. an EMS AI helping you to save someone's life in an emergency.

3. an AI coach that can teach you a new sport or activity.

4. an omnipresent domain-expert that can show you how to make a gourmet meal, repair an engine, or perform a traditional tea ceremony.

5. a personal assistant that can anticipate your information need (what's that person's name? where's the exit? who's the most interesting person here? etc.) and whisper the answer in your ear just as you need it.

Now, add all of the above to an AR capability where you can now think or speak of something interesting and complex, and have it visualized right before your eyes. With this capability, I could augment my imagination with almost super-human capabilities that allow one to solve complex problems almost as if it was an internal mental monologue.

All of these scenarios are just a short hop from where were at now, so mark my words: we will have "borgs" like those described above long before we reach anything like general AI.

Re: Dall-E 2

#96

Very cool stuff. For me, the most interesting was the ability to take a piece of art and generate variations of it. Have a favorite painter? Here's 10,000 new paintings like theirs.

Well, one of my favorite painters is Henri Rousseau, and one of his great paintings is War, 1984: https://www.henrirousseau.net/war.jsp However, this painting has themes of violence and politics plus some nude dead bodies, so it violates the content policy: "Our content policy does not allow users to generate violent, adult, or political content, among other categories." So what you'd get is some kind of sanitized wa…

They are being rightly cautious. It’s going to take time to figure out good practice with these tools. Everyone calling out basic caution as “dystopian” is really over the top.

I’ve been using tools like this for over a year now. Even with filtered dataset and filtered interface, they can make images that would make the Fangoria crowd blush if you put the slightest effort into it.

It’s one thing to be able to make brain-wrenching images with a lot of photoshop effort (or digging hard enough in the dark corners of the internet). It’s another thing entirely give anyone the ability to spew out thousands of them trivially.

Re: Dall-E 2

#97
post #90
post #73

Earlier quoted context omitted.

What always bother me with this stuff is, well, you say one approach is more sensible than the other because the images happen to come out more pleasing. But there's no real rhyme or reason, it is a sort of alchemy. Is text encoding strictly worse or is it an artifact of the implementation? And if it is strictly worse, which is probably the case, why specifically? What is actually going on here? I can't argue that th…

Yeah, I mean you're right that ultimately the proof is in the pudding. But I do think we could have guessed that this sort of approach would be better (at least at a high level - I'm not claiming I could have predicted all the technical details!). The previous approaches were sort of the best that people could do without access to the training data and resources - you had a pretrained CLIP encoder that could tell you…

I don't think it is actually painting at all but I need to read the paper carefully.

I think it is using a free text query to select the best possible clipart from a big library and blends it together. Still very interesting and useful.

It would be extremely impressive if the "Kuala dunking a basketball" had a puddle on the court in which it was reflected correctly, that would be mind blowing.

Re: Dall-E 2

#98
post #78

Earlier quoted context omitted.

Quoted post unavailable.

> this service should be provided to me and if it isn't done how I want it that's infringing on me somehow That is an extremely uncharitable interpretation of: > I wish I could have a version with the training wheels taken off.

I would have responded differently had that been the statement. But many of the responses were more than that.

Re: Dall-E 2

#99
post #82
post #59

I'm only part way through the paper, but what struck me as interesting so far is this: In other text-to-image algorithms I'm familiar with (the ones you'll typically see passed around as colab notebooks that people post outputs from on Twitter), the basic idea is to encode the text, and then try to make an image that maximally matches that text encoding. But this maximization often leads to artifacts - if you ask for…

This isn't something I'm knowledgeable on so forgive my simplification but is this like a sort of micro services for AI. Each AI takes their turn handing some aspect, another sort of mediates among them?

I'd say Dall-E 2 is a little more unified - they do have multiple networks, but they're trained to work together. The previous approaches I was talking about are a lot more like the microservices analogy. Someone published a model (called CLIP) that can say "how much does this image look like a sunset". Someone else published a totally different model (e.g. VQGAN) that can generate images (but with no way to provide text prompts). A third person figures out a clever way to link the two up - have the VQGAN make an image, ask CLIP how much it looks like a sunset, and use backpropagation to adjust the image a little, repeat until you have a sunset. Each component is it's own thing, and VQGAN and CLIP don't know anything about one another.

Re: Dall-E 2

#100
post #6
post #3

Preventing Harmful Generations We’ve limited the ability for DALL·E 2 to generate violent, hate, or adult images. By removing the most explicit content from the training data, we minimized DALL·E 2’s exposure to these concepts. We also used advanced techniques to prevent photorealistic generations of real individuals’ faces, including those of public figures. "And we've also closed off a huge range of potentially int…

I've been on HN for years and I still can't figure out how to format text as a quote I don't think there is a way comparable to markdown, since the formatting options are limited: https://news.ycombinator.com/formatdoc So your options are literal quotes, "code" formatting like you've done, italics like I've done, or the '>' convention, but that doesn't actually apply formatting. Would be nice if it were added.

And the "code" formatting for quotes is generally a bad choice because people read on a variety of screen sizes, and "code" formatting can screw that up (try reading the quote with a really narrow window).
Post reply on HN