Live data from Hacker News

AI World Clocks

clocks.brianmoore.com

191–200 of 404 posts

Re: AI World Clocks

#192

Watching this over the past few minutes, it looks like Kimi K2 generates the best clock face most consistently. I'd never heard of that model before today! Qwen 2.5's clocks, on the other hand, look like they never make it out of the womb.

I find that Kimi K2 looks the best, but i've noticed the time is often wrong!

Re: AI World Clocks

#193

Watching this over the past few minutes, it looks like Kimi K2 generates the best clock face most consistently. I'd never heard of that model before today! Qwen 2.5's clocks, on the other hand, look like they never make it out of the womb.

Qwen's clocks are highly entertaining. Like if you asked an alien "make me a clock".

Re: AI World Clocks

#194
post #104

Earlier quoted context omitted.

It could be that the prompt is accidentally (or purposefully) more optimised for Kimi K2, or that Kimi K2 is better trained on this particular data. LLM's need "prompt engineers" for a reason to get the most out of a particular model.

It's not fair to use prompts tailored to a particular model when doing comparisons like this - one shot results that generalize across a domain demonstrate solid knowledge of the domain. You can use prompting and context hacking to get any particular model to behave pseudo-competently in almost any domain, even the tiny This experiment, however, clearly states the goal with this prompt: `Create HTML/CSS of an analog…

[deleted]

Re: AI World Clocks

#195
Reminds me of the Alzheimer's "draw a clock" test.

Makes me think that LLMs are like people with dementia! Perhaps it's the best way to relate to an LLM?

Re: AI World Clocks

#196
post #104

Earlier quoted context omitted.

It could be that the prompt is accidentally (or purposefully) more optimised for Kimi K2, or that Kimi K2 is better trained on this particular data. LLM's need "prompt engineers" for a reason to get the most out of a particular model.

How much engineering do prompt engineers do? Is it engineering when you add "photorealistic. correct number of fingers and teeth. High quality." to the end of a prompt? we should call them "prompt witch doctors" or maybe "prompt alchemists".

"How is engineering a real science? You just build the bridge so it doesn't fall down."

Re: AI World Clocks

#199

Earlier quoted context omitted.

I've noticed that image models are particularly bad at modifying popular concepts in novel ways (way worse "generalization" than what I observe in language models).

Maybe LLMs always fail to generalize outside their data set, and it’s just less noticeable with written language.

This is it. They’re language models which predict next tokens probabilistically and a sampler picks one according to the desired ”temperature”. Any generalization outside their data set is an artifact of random sampling: happenstance and circumstance, not genuine substance.

Re: AI World Clocks

#200

hi, I made this. thank you for posting. I love clocks and I love finding the edges of what any given technology is capable of. I've watched this for many hours and Kimi frequently gets the most accurate clock but also the least variation and is most boring. Qwen is often times the most insane and makes me laugh. Which one is "better?"

Why is this different per user? I sent this to a few friends and they all see different things from what i'm seeing, for the same time..?
Post reply on HN