Live data from Hacker News

AI World Clocks

clocks.brianmoore.com

301–310 of 404 posts

Re: AI World Clocks

#303
post #104

Earlier quoted context omitted.

It could be that the prompt is accidentally (or purposefully) more optimised for Kimi K2, or that Kimi K2 is better trained on this particular data. LLM's need "prompt engineers" for a reason to get the most out of a particular model.

How much engineering do prompt engineers do? Is it engineering when you add "photorealistic. correct number of fingers and teeth. High quality." to the end of a prompt? we should call them "prompt witch doctors" or maybe "prompt alchemists".

> we should call them "prompt witch doctors" or maybe "prompt alchemists".

Oh absolutely not! Only in engineering you are allowed to get called an engineer for no apparent reason, do that in other white collar and you are behind the bars because of fraudulent claims.

Re: AI World Clocks

#306

Non-determinism at it's finest. The clock is perfect, the refresh happens, the clock looks like a Dali painting.

Last year I wrote a simple system using Semantic Kernel, backed by functions inside Microsoft Orleans, which for the most part was a business logic DSL processor by LLM. Your business logic was just text, and you gave it the operation as text.

Nothing could be relied upon to be deterministic, it was so funny to see it try to do operations.

Recently I re-ran it with newer models and was drastically better, especially with temperature tweaks.

Re: AI World Clocks

#307
post #82

Amazing, some people are so enamored with LLMs who use them for soft outcomes, and disagree with me when I say be careful they're not perfect -- this is such a great non technical way to explain the reality I'm seeing when using on hard outcome coding/logic tasks. "Hey this test is failing", LLM deletes test , "FIXED!"

Something that struck me when I was looking at the clocks is that we know what a clock is supposed to look and act like. What about when we don't know what it's supposed to look like? Lately I've been wrestling with the fact that unlike, say, a generalized linear model fit to data with some inferential theory, we don't have a theory or model for the uncertainty about LLM products. We recognize when it's off about thi…

I need to be delicate with wording here, but this is why it's a worry that all the least intelligent people you know could be using AI.

It's why non-coders think it's doing an amazing job at software.

But it's worryingly why using it for research, where you necessarily don't know what you don't know, is going to trip up even smarter people.

Re: AI World Clocks

#308
I wonder which model will silently be updated and suddenly start drawing clocks with Audemars-Piguet-level kind of complications.

Re: AI World Clocks

#309

hi, I made this. thank you for posting. I love clocks and I love finding the edges of what any given technology is capable of. I've watched this for many hours and Kimi frequently gets the most accurate clock but also the least variation and is most boring. Qwen is often times the most insane and makes me laugh. Which one is "better?"

Clock drawing is widely used as a test for assessing dementia. Sometimes the LLMs fail in ways that are fairly predictable if you're familiar with CSS and typical shortcomings of LLMs, but sometimes they fail in ways that are less obvious from a technical perspective but are exactly the same failure modes as cognitively-impaired humans. I think you might have stumbled upon something surprisingly profound. https://www…

> Clock drawing is widely used as a test for assessing dementia

Interestingly, clocks are also an easy tell for when you're dreaming, if you're a lucid dreamer; they never work normally in dreams.

Post reply on HN