Live data from Hacker News

Agentic pelican on a bicycle

robert-glaser.de

41–50 of 77 posts

Re: Agentic pelican on a bicycle

#41
post #29

Could this be improved if the evaluation was done by an independent sub-agent?

Is it running out of space in its context window?

My rational is that perhaps it's being biased towards continuing doing what it's doing, or biased towards telling that it has done a good job and not being self-critical.

Re: Agentic pelican on a bicycle

#42
post #36

Earlier quoted context omitted.

The technical definition of an agent is an LLM being called in a loop, some of which calls include tool definitions. That's exactly what this is.

"Effectful loops" or "augmented loops" are much more descriptive of what is actually going on & do not confuse the reader w/ incoherent definitions of "agency".

So this whole thread was just you trying to express that you don't like the common accepted definition of the term "agent"?

Re: Agentic pelican on a bicycle

#43
I love this experiment and am surprised that the Claude models performed that much better than the competition. Opus was particularly impressive both in the quality itself and the ability to iterate meaningfully.

Now... Was this article LLM written?

This part triggered all my LLM flags: ``` Adding a bicycle chain isn’t just decoration—it shows understanding of mechanical relationships. The wheel spokes, the adjusted proportions—these are signs of vision-driven refinement working as intended. ```

Re: Agentic pelican on a bicycle

#44
post #42

Earlier quoted context omitted.

"Effectful loops" or "augmented loops" are much more descriptive of what is actually going on & do not confuse the reader w/ incoherent definitions of "agency".

So this whole thread was just you trying to express that you don't like the common accepted definition of the term "agent"?

I prefer coherent definitions instead of corporate marketing. Whether I like the term or not is secondary. Judging by past instances of this same phenomenon I expect the word to lose all meaning as more companies start telling their customers about their "agentic" offerings.

Re: Agentic pelican on a bicycle

#45
post #8

What I take from this is that LLMs are somewhat miraculous in generation but terrible at revision. Especially with images, they are very resistant to adjusting initial approaches. I wonder if there is a consistent way to force structural revisions. I have found Nano Banana particularly terrible at revisions, even something like "change the image dimensions to..." it will confidently claim success but do nothing.

I see this all the time when asking Claude or ChapGPT to produce a single-page two-column PDF summarizing the conclusions of our chat. Literally 99% of the time I get a multi-page unpredictably-formatted mess, even after gently asking over and over for specific fixes to the formatting mistake/s. And as you say, they cheerfully assert that they've done the job, for real this time, every time.

Ask for the asciidoc and asciidoctor command to make a PDF instead. Chat bots aren’t designed to make PDFs. They are just trying to use tools in the background, probably starting with markdown.

Re: Agentic pelican on a bicycle

#46

Earlier quoted context omitted.

It's going to become the "MP3 sizzle" that young people at the time started to prefer once compressed audio became the norm on iPods and other portable music players, along with film grain and the judder of 24fps video. Artifacts imposed by the medium themselves become desirable once they become normal an associated and in fact signs of "quality", when, in fact, they are introduced noise and distortion to an otherwis…

With vinyl warmth is the result of a deliberate process. Professional masters are done specifically for vinyl to accommodate its quirks which truly changes the sound. They have to clamp down the dynamic range and tidy low frequencies or the needle will skip. Recordings with lots of busy high frequency information also can’t be physically captured properly in the cut. The resulting master is a version that purposefull…

Same vein https://open.substack.com/pub/animationobsessive/p/the-toy-s...

Pixar films were setup with the idea of being put on film so the DVD digital transfers color is all wrong.

Re: Agentic pelican on a bicycle

#47

I love this experiment and am surprised that the Claude models performed that much better than the competition. Opus was particularly impressive both in the quality itself and the ability to iterate meaningfully. Now... Was this article LLM written? This part triggered all my LLM flags: ``` Adding a bicycle chain isn’t just decoration—it shows understanding of mechanical relationships. The wheel spokes, the adjusted…

I mean at some point you have to evaluate the content on its merit and they have a point — a chain is functional not just decorative in its precise placement.

Re: Agentic pelican on a bicycle

#48
post #47

I love this experiment and am surprised that the Claude models performed that much better than the competition. Opus was particularly impressive both in the quality itself and the ability to iterate meaningfully. Now... Was this article LLM written? This part triggered all my LLM flags: ``` Adding a bicycle chain isn’t just decoration—it shows understanding of mechanical relationships. The wheel spokes, the adjusted…

I mean at some point you have to evaluate the content on its merit and they have a point — a chain is functional not just decorative in its precise placement.

That phrase template isn’t just overdone—it's something some text models are obsessed with. The em-dashes, the contrastive language—these are signs of LLMs being asked to summarize or expand a compelling blog post.

Re: Agentic pelican on a bicycle

#49
post #36

Earlier quoted context omitted.

The technical definition of an agent is an LLM being called in a loop, some of which calls include tool definitions. That's exactly what this is.

"Effectful loops" or "augmented loops" are much more descriptive of what is actually going on & do not confuse the reader w/ incoherent definitions of "agency".

[deleted]

Re: Agentic pelican on a bicycle

#50

What I take from this is that LLMs are somewhat miraculous in generation but terrible at revision. Especially with images, they are very resistant to adjusting initial approaches. I wonder if there is a consistent way to force structural revisions. I have found Nano Banana particularly terrible at revisions, even something like "change the image dimensions to..." it will confidently claim success but do nothing.

A thing I've been noticing across the board is that current generative AI systems are horrible at composition. It’s most obvious in image generation models where the composition and blocking tend to be jarringly simple and on point (hyper-symmetry, all-middleground, or one of like three canned "artistic" compositions) no matter how you prompt them, but you see it in things like text output as well once you notice it.

I suspect this is either a training data issue, or an issue with the people building these things not recognizing the problem, but it's weird how persistent and cross-model the issue is, even in model releases that specifically call out better/more steerable composition behavior.

Post reply on HN