Earlier quoted context omitted.
> Google's Gemini LLM model was not to blame for the image generation weirdness. That would be like blaming DALL-E weirdness on GPT-4. The way I read the Gemini technical report, it seemed like, unlike GPT-4 vs DALL-E, Gemini was pretrained with multimodal outputs. Is that not the case?
Is that right? I didn't think Gemini was generating images directly, I assumed it was using a separate image generation tool. The paper here https://arxiv.org/pdf/2403.05530.pdf has a model card for Gemini 1.5 Pro that says: Output(s): Generated text in response to the input (e.g., an answer to the question, a summary of multiple documents, comparing documents/videos).
That feels like it runs counter to this statement from the Gemini 1.0 technical report[0]:
> Gemini models are trained to accommodate textual input interleaved with a wide variety of audio and visual inputs, such as natural images, charts, screenshots, PDFs, and videos, and they can produce text and image outputs