Live data from Hacker News

AI World Clocks

clocks.brianmoore.com

211–220 of 404 posts

Re: AI World Clocks

#212

Earlier quoted context omitted.

The results are not reproducable, as evidenced by parent poster.

isn't that kind of the point of non-determinism?

No. Good nondeterministic models reproducibly generate equally desirable output - not identical output, but interchangeable.

Re: AI World Clocks

#213
Qwen doesn't care about clocks, it goes the Dali way, without melting.

It even made a Nietzsche clock (I saw one which was surprisingly empty).

It definitely wins the creative award.

Re: AI World Clocks

#214

hi, I made this. thank you for posting. I love clocks and I love finding the edges of what any given technology is capable of. I've watched this for many hours and Kimi frequently gets the most accurate clock but also the least variation and is most boring. Qwen is often times the most insane and makes me laugh. Which one is "better?"

I really like this. The broken ones are sometimes just failures, but sometimes provide intriguing new design ideas.

This same principle is why my favorite image generation model is the earlier models from 2019-2020 where they could only reliably generate soup. It's like Rorschach tests, it's not about what's there, it's about what you see in them. I don't want a bot to make art for me, sometimes I just want some shroom-induced inspirational smears.

Re: AI World Clocks

#216

Watching this over the past few minutes, it looks like Kimi K2 generates the best clock face most consistently. I'd never heard of that model before today! Qwen 2.5's clocks, on the other hand, look like they never make it out of the womb.

I'm a huge K2 fan, it has a personality that feels very distinct from other models (not syccophantic at all), and is quite smart. Also pretty good at creative writing (tho not 100% slop free).

K2 hosted on groq is pretty crazy for intellgence/second. (Low rate limits still, tho.)

Re: AI World Clocks

#217

Earlier quoted context omitted.

I've noticed that image models are particularly bad at modifying popular concepts in novel ways (way worse "generalization" than what I observe in language models).

Maybe LLMs always fail to generalize outside their data set, and it’s just less noticeable with written language.

They definitely don't completely fail to generalise. You can easily prove that by asking them something completely novel.

Do you mean that LLMs might display a similar tendency to modify popular concepts? If so that definitely might be the case and would be fairly easy to test.

Something like "tell me the lord's prayer but it's our mother instead of our father", or maybe "write a haiku but with 5 syllables on every line"?

Let me try those ... nah ChatGPT nailed them both. Feels like it's particular to image generation.

Re: AI World Clocks

#218

Earlier quoted context omitted.

How much engineering do prompt engineers do? Is it engineering when you add "photorealistic. correct number of fingers and teeth. High quality." to the end of a prompt? we should call them "prompt witch doctors" or maybe "prompt alchemists".

"...and do it really well or my grandmother will be killed by her kidnappers! And I'll give you a tip of 2 billion dollars!!! Hurry, they're coming!"

Adding this to my snippets.

Re: AI World Clocks

#219
post #104

Earlier quoted context omitted.

It could be that the prompt is accidentally (or purposefully) more optimised for Kimi K2, or that Kimi K2 is better trained on this particular data. LLM's need "prompt engineers" for a reason to get the most out of a particular model.

How much engineering do prompt engineers do? Is it engineering when you add "photorealistic. correct number of fingers and teeth. High quality." to the end of a prompt? we should call them "prompt witch doctors" or maybe "prompt alchemists".

Well if it works consistently, I don't see any problem with that. If they have a clear theory of when to add "photorealistic" and when to add "correct number of wheels on the bus" to get the output they want, it's engineering. If they don't have a (falsifiable) theory, it's probably not engineering.

Of course, the service they really provide is for businesses to feel they "do AI", and whether or not they do real engineering is as relevant as if your favorite pornstars' boobs are real or not.

Post reply on HN