Live data from Hacker News

The Generative Burrito Test

generativist.com

31–40 of 56 posts

Re: The Generative Burrito Test

#31
post #8

Oh wow, I've been hearing about Nano Banana Pro in random stuff lately, but as a layman the difference is stark. It's the only one that actually looks like a partially eaten burrito at all to me. The others all look like staged marketing fake food, if I'm being generous (only a few actually approach that, most just look wrong).

The NBP looks like a mock of food to me - the unwrapped burrito on a single piece of intact tinfoil, a table where the grain goes all wonky, an almost pastry looking tortilla, hyperrealistic beans and there's something wrong with the focal plane. It's just not as plasticy and oversaturated as the others.

Hyperrealistic beans? The focal plane? You are reaching really hard here.

The table grain is the only thing that gives it away - if it weren't for that no one without advance warning is going to notice that it's not real.

Re: The Generative Burrito Test

#32

An interesting American culinary divide is between Scottsdale and Phoenix homemade burritos. The former being close to the Midwest variety, the latter to a Sonoran style. Even ignoring the Heinz bean outliers, these are all decidedly Scottsdale. With one exception. All hail Nano Banana.

Just convert a tamale extruder to take in raw tortilla dough and bake the whole thing at once to cook the tortilla around the fillings.

Re: The Generative Burrito Test

#33
post #27

With llms there is a secondary training step to turn a foundational model into a chat bot. Is these something similar going on with these image generation models, that is making them all tend towards making pretty clean images and stopping them making half eaten food even if they have the capabilities?

In terms of prompt adherence, there are two issues with most image generation models, neither of which apply to Nano Banana:

1. The text encoders are primitive (e.g. CLIP) and have difficulty with nuance, such as "partially eaten", and model training can only partially overcome it. It's the same issue with the now-obsolete "half-filled" wine glass test.

2. Most models are diffusion-based, which means it denoises the entire image simultaneously. If it fails to account for the nuance in the first few passes, it can't go back and fix it.

I believe some image generation AIs were RLHFed like chat bot LLMs, but moreso to improve aesthetics rather than prompt adherence.

Re: The Generative Burrito Test

#34

One of my tests for new image generation models is professional food photography, particularly in cases where the food has constraints, such as "a peanut butter and jelly sandwich in the shape of a Rubik’s cube" (blog post from 2022 for DALL-E 2: https://minimaxir.com/2022/07/food-photography-ai/ ) For some reason ever since DALL-E 2, all food models seem to generate obviously fake food and/or misinterpret the fun co…

I tried having Claude generate a prompt for Seedream and got this: https://imgur.com/a/6xX5TDE

I can kind of see what you mean in that it went for realism in the aesthetics, but not the object... but that last one would probably fool me if I was scrolling

Re: The Generative Burrito Test

#35

One of my tests for new image generation models is professional food photography, particularly in cases where the food has constraints, such as "a peanut butter and jelly sandwich in the shape of a Rubik’s cube" (blog post from 2022 for DALL-E 2: https://minimaxir.com/2022/07/food-photography-ai/ ) For some reason ever since DALL-E 2, all food models seem to generate obviously fake food and/or misinterpret the fun co…

Nano-Banana does a (inter)stellar job with food based prompts.

https://mordenstar.com/portfolio/wontauns

Re: The Generative Burrito Test

#36
post #8

Oh wow, I've been hearing about Nano Banana Pro in random stuff lately, but as a layman the difference is stark. It's the only one that actually looks like a partially eaten burrito at all to me. The others all look like staged marketing fake food, if I'm being generous (only a few actually approach that, most just look wrong).

This shows some gaps in the "same prompt to every model" approach to benchmarking models. I get that it's allows ensuring you're testing the model capabilities vs prompts, but most models are being post-trained with very different formats of prompting. I use Seedream in production so I was a little suspicious of the gap: I passed Bytedance's official prompting guide, OPs prompt, and your feedback to Claude Opus 4.5 a…

not adhering to the prompt guide is def a valid strong criticism. resampling i think less so for the demo just because fewer people look at k samples per model, so just taking literally the first one has the fewest of my own biases injected into it

Re: The Generative Burrito Test

#37

One of my tests for new image generation models is professional food photography, particularly in cases where the food has constraints, such as "a peanut butter and jelly sandwich in the shape of a Rubik’s cube" (blog post from 2022 for DALL-E 2: https://minimaxir.com/2022/07/food-photography-ai/ ) For some reason ever since DALL-E 2, all food models seem to generate obviously fake food and/or misinterpret the fun co…

I tried having Claude generate a prompt for Seedream and got this: https://imgur.com/a/6xX5TDE I can kind of see what you mean in that it went for realism in the aesthetics, but not the object... but that last one would probably fool me if I was scrolling

Those are better than usual: I've gotten generations from earlier models that are just a normal colorful Rubix's cube between two slices of bread.

Re: The Generative Burrito Test

#38
post #8

Oh wow, I've been hearing about Nano Banana Pro in random stuff lately, but as a layman the difference is stark. It's the only one that actually looks like a partially eaten burrito at all to me. The others all look like staged marketing fake food, if I'm being generous (only a few actually approach that, most just look wrong).

This shows some gaps in the "same prompt to every model" approach to benchmarking models. I get that it's allows ensuring you're testing the model capabilities vs prompts, but most models are being post-trained with very different formats of prompting. I use Seedream in production so I was a little suspicious of the gap: I passed Bytedance's official prompting guide, OPs prompt, and your feedback to Claude Opus 4.5 a…

100%. Between tuning prompt variations depending on the model and allowing a minimum number of re-rolls, this is why it takes a while to publish results from the newest models on my GenAI comparison site.

Including a "total rolls" is a very valuable metric since it helps indicate how steerable the model is.

Re: The Generative Burrito Test

#39

Earlier quoted context omitted.

This shows some gaps in the "same prompt to every model" approach to benchmarking models. I get that it's allows ensuring you're testing the model capabilities vs prompts, but most models are being post-trained with very different formats of prompting. I use Seedream in production so I was a little suspicious of the gap: I passed Bytedance's official prompting guide, OPs prompt, and your feedback to Claude Opus 4.5 a…

not adhering to the prompt guide is def a valid strong criticism. resampling i think less so for the demo just because fewer people look at k samples per model, so just taking literally the first one has the fewest of my own biases injected into it

I actually think it's ok to inject your own bias here: if you're deploying these models in production, then you probably test on your own domain other than half eaten burritos lol

But individual users usually iterate/pick, so just sharing a blurb about your preference is probably enough if you choose 1 of n

Re: The Generative Burrito Test

#40

Earlier quoted context omitted.

The NBP looks like a mock of food to me - the unwrapped burrito on a single piece of intact tinfoil, a table where the grain goes all wonky, an almost pastry looking tortilla, hyperrealistic beans and there's something wrong with the focal plane. It's just not as plasticy and oversaturated as the others.

Hyperrealistic beans? The focal plane? You are reaching really hard here. The table grain is the only thing that gives it away - if it weren't for that no one without advance warning is going to notice that it's not real.

I am a huge AI skeptic, check my comment history.

I agree with you. The Nano Banana Pro burrito is almost perfect, the wood grain direction/perspective is the only questionable element.

Almost no one would ID that as being AI.

Post reply on HN