Live data from Hacker News

AI World Clocks

clocks.brianmoore.com

401–404 of 404 posts

Re: AI World Clocks

#401
post #82

Amazing, some people are so enamored with LLMs who use them for soft outcomes, and disagree with me when I say be careful they're not perfect -- this is such a great non technical way to explain the reality I'm seeing when using on hard outcome coding/logic tasks. "Hey this test is failing", LLM deletes test , "FIXED!"

Something that struck me when I was looking at the clocks is that we know what a clock is supposed to look and act like. What about when we don't know what it's supposed to look like? Lately I've been wrestling with the fact that unlike, say, a generalized linear model fit to data with some inferential theory, we don't have a theory or model for the uncertainty about LLM products. We recognize when it's off about thi…

I built an ML classifier for product categories way back, as I added more classes/product types, individual class PR metrics improved--I kept adding more and more until I ended up with ~2,000 classes.

My intuition is at the start when I was like "choose one of these 10 or unknown", that unknown left a big gray area, so as I added more classes the model could say "I know it's not X, because it's more similar to Y"

I feel like in this case though, the broken clocks are broken because they don't serve the purpose of visually transmitting information, they do look like clocks tho. I'm sure if you fed the output back into the LLM and ask what time it is it would say IDK, or more likely make something up and be wrong. (at least the egregious ones where the hands are flying everywhere)

Re: AI World Clocks

#402

Try adding to the prompt that it has a PhD in Computer Science and have many methods for dealing with complexity. This gives better results, at least for me.

Why does that give better results? Is this phenomena measurable? How would "you have a phd in computer science" change its ability to interpret prose? Every interaction with an LLM seems like superstition.

Because ie. a forum thread that contains this often have better answers, and LLMs are trained on data from the Internet. It's just statistics.

Re: AI World Clocks

#403

Earlier quoted context omitted.

Also the amount of adjacent remarks being always topical flsvor confusion is cartoonish. Im playing with ideas for making thst better

You're absolutely right! People should pay attention to this broadly applicable and important consideration.

I have some ideas for a mini logic solver mcp thst might help. Currently making it sortah simulate using such a tool

Its one of those things where it feels like itd be easy to get copycats even if theres a market

Re: AI World Clocks

#404
Wonderful. I don’t particularly care if it is or is not a valid test. I like the “wrong” renderings better. Some are hilarious, some … inspired.
Post reply on HN