Live data from Hacker News

AI World Clocks

clocks.brianmoore.com

231–240 of 404 posts

Re: AI World Clocks

#231
post #104

Watching this over the past few minutes, it looks like Kimi K2 generates the best clock face most consistently. I'd never heard of that model before today! Qwen 2.5's clocks, on the other hand, look like they never make it out of the womb.

It could be that the prompt is accidentally (or purposefully) more optimised for Kimi K2, or that Kimi K2 is better trained on this particular data. LLM's need "prompt engineers" for a reason to get the most out of a particular model.

I think the selection of models is a bit off. Haiku instead of Sonnet for example. Kimi K2's capabilities are closer to Sonnet than to Haiku. GPT-5 might be in the non-reasoning mode, which routes to a smaller model.

Re: AI World Clocks

#232

I'm having a hard time believing this site is honest, especially with how ridiculous the scaling and rotation of numbers is for most of them. I dumped his prompt into chatgpt to try it myself and it did create a very neat clock face with the numbers at the correct position+animated second hand, it just got the exact time wrong, being a few hours off. Edit: the time may actually have been perfect now that I account fo…

i read that the OP limited the output to 2000 tokens.

Re: AI World Clocks

#233

Earlier quoted context omitted.

isn't that kind of the point of non-determinism?

No. Good nondeterministic models reproducibly generate equally desirable output - not identical output, but interchangeable.

oh I see, thank you for clarifying

Re: AI World Clocks

#234

I'm having a hard time believing this site is honest, especially with how ridiculous the scaling and rotation of numbers is for most of them. I dumped his prompt into chatgpt to try it myself and it did create a very neat clock face with the numbers at the correct position+animated second hand, it just got the exact time wrong, being a few hours off. Edit: the time may actually have been perfect now that I account fo…

i read that the OP limited the output to 2000 tokens.

^ this! there's a lot of clocks to generate so I've challenged it to stick to a small(er) amount of code

Re: AI World Clocks

#235

Earlier quoted context omitted.

"...and do it really well or my grandmother will be killed by her kidnappers! And I'll give you a tip of 2 billion dollars!!! Hurry, they're coming!"

Ive heard this actually works annoyingly well

We've created technology so sophisticated it is vulnerable to social engineering attacks.

Re: AI World Clocks

#236

Since the first (good) image generation models became available, I've been trying to get them to generate an image of a clock with 13 instead of the usual 12 hour divisions. I have not been successful. Usually they will just replace the "12" with a "13" and/or mess up the clock face in some other way. I'd be interested if anyone else is successful. Share how you did it!

I gave this "riddle" to various models:

> The farmer and the goat are going to the river. They look into the sky and see three clouds shaped like: a wolf, a cabbage and a boat that can carry the farmer and one item. How can they safely cross the river?

Most of them are just giving the result to the well known river crossing riddle. Some "feel" that something is off, but still have a hard time to figure out that wolf, boat and cabbage are just clouds.

Re: AI World Clocks

#238
post #219

Earlier quoted context omitted.

Well if it works consistently, I don't see any problem with that. If they have a clear theory of when to add "photorealistic" and when to add "correct number of wheels on the bus" to get the output they want, it's engineering. If they don't have a (falsifiable) theory, it's probably not engineering. Of course, the service they really provide is for businesses to feel they "do AI", and whether or not they do real engi…

Maybe we could keep the conversation out of the gutter.

Porn is taxable income, not the gutter.

Re: AI World Clocks

#239
post #82

Amazing, some people are so enamored with LLMs who use them for soft outcomes, and disagree with me when I say be careful they're not perfect -- this is such a great non technical way to explain the reality I'm seeing when using on hard outcome coding/logic tasks. "Hey this test is failing", LLM deletes test , "FIXED!"

Something that struck me when I was looking at the clocks is that we know what a clock is supposed to look and act like.

What about when we don't know what it's supposed to look like?

Lately I've been wrestling with the fact that unlike, say, a generalized linear model fit to data with some inferential theory, we don't have a theory or model for the uncertainty about LLM products. We recognize when it's off about things we know are off, but don't have a way to estimate when it's off other than to check it against reality, which is probably the exception to how it's used rather than the rule.

Re: AI World Clocks

#240
post #208

Earlier quoted context omitted.

"How is engineering a real science? You just build the bridge so it doesn't fall down."

Nah. Actual engineers have professional standards bodies and legal liability when they shirk and the bridge falls down or the plane crashes or your wiring starts on fire. Software "engineers" are none of those things but can at least emulate the approaches and strive for reproducibility and testability. Skilled craftsman; not engineers. Prompt "engineers" is yet another few steps down the ladder, working out mostly b…

  The battle on the use of language around engineer has long been lost
That's really the core of the issue: We're just having the age-old battle of prescriptivism vs descriptivism again. An "engineer", etymologically, is basically just "a person who comes up with stuff", one who is "ingenious". I'm tempted to say it's you prescriptivists who are making a "battle" out of this.

  subjective creative exercise of writing prompts
Implying that there are no testable results, no objective success or failure states? Come on man.
Post reply on HN