Live data from Hacker News

Using “underdrawings” for accurate text and numbers

samcollins.blog

31–40 of 140 posts

Re: Using “underdrawings” for accurate text and numbers

#33

inb4 this technique is subsumed into the next MoE model release LLMs are evolving so fast I wouldn’t be surprised if this technique was not needed in <6 months

I don't think the MoE part has anything to do with it, but the current gen of multimoddal models can do thinking interleaved with autoregressive(?*) image-gen so it's probably not long before they bake this into the RL process, same way native thought obviated need for "think carefully step by step" prompts.

Re: Using “underdrawings” for accurate text and numbers

#35
post #30
post #25

Earlier quoted context omitted.

> due to fundamental limitations People keep throwing this phrase around in relation to LLMs, when not a single “fundamental limitation” has been rigorously demonstrated to exist, and many tasks that were claimed to be impossible for LLMs two years ago supposedly due to “fundamental limitations” (e.g. character counting or phonetics) are non-issues for them today even without tools.

of course, if you choose to ignore all the limitations they indeed have no limitations.

Nobody says they have no limitations. The question is are those limitation fundamental, i.e. can we expect improvement, say within a year.

Re: Using “underdrawings” for accurate text and numbers

#36
post #29

The standard objection: if the LLM is supposedly intelligent, why can’t it figure out on its own that this two-step process would achieve a better result?

Part of the problem is that it isn't the LLM making the image directly itself, it's the LLM repeatedly prompting edits for a separate edit diffusion model. The Gemini reasoning summary shows part of this. The style of some of the images makes it also clear that it uses an Imagen 4 derived diffusion model underneath.

Re: Using “underdrawings” for accurate text and numbers

#38
post #25

I'm glad that we're making progress towards a deeper understanding of what LLMs are inherently good at and what they're inherently bad at (not to say incapable of doing, but stuff that is less likely to work due to fundamental limitations). There's similarity here with, for example, defining the architecture of software, but letting an LLM write the functions. Or asking an LLM to write you the SQL query for your data…

> due to fundamental limitations People keep throwing this phrase around in relation to LLMs, when not a single “fundamental limitation” has been rigorously demonstrated to exist, and many tasks that were claimed to be impossible for LLMs two years ago supposedly due to “fundamental limitations” (e.g. character counting or phonetics) are non-issues for them today even without tools.

This is kind of my point, we need to get better at describing the limitations and study them. It seems extremely clear that there are limitations, and not just temporary ones, but structural limitations that existed at the beginning and continue to persist.

Re: Using “underdrawings” for accurate text and numbers

#39
post #29

The standard objection: if the LLM is supposedly intelligent, why can’t it figure out on its own that this two-step process would achieve a better result?

[flagged]

Every decent human artist knows to draw a sketch before painting something.

Re: Using “underdrawings” for accurate text and numbers

#40
post #30

Earlier quoted context omitted.

of course, if you choose to ignore all the limitations they indeed have no limitations.

Nobody says they have no limitations. The question is are those limitation fundamental, i.e. can we expect improvement, say within a year.

When I talk about fundamental limitations, I mean limitations that can't be solved, even if they could be improved.

We have improved hallucinations significantly, and yet it seems clear that they are inherent to the technology and so will always exist to some extent.

Post reply on HN