Live data from Hacker News

Choosing an AI model: one prompt, 11 models, different results

netlify.com

101–105 of 105 posts

Re: Choosing an AI model: one prompt, 11 models, different results

#101
post #57

Earlier quoted context omitted.

That's probably because the actual coding benchmarks were saturated several years ago.

Which is why you should perform your own benchmarks against your own software stack.

Agree. To this I would add, many things are saturated even for smaller models, which tend to be much cheaper and faster.

On many of my tests, there was no difference in the result between the smaller and bigger model, but there was a big difference in speed and price.

Re: Choosing an AI model: one prompt, 11 models, different results

#102

> Build a one-page site for a neighbourhood coffee shop: opening hours, the address, a short menu and a photo. Nothing on it changes unless I edit it myself. If that's the entire prompt, it's quite depressing how much alike these all look. I appreciate some of the details from the Opus 5 version, but I can't help but strongly feel the AI vibes emanating from that design.

Don't blame LLMs if the zeitgeist for coffee shop websites is sepia tones and a Papyrus-like font.

My university's Principles of WebDev course 10 years ago had a similar assignment and the results all ended up looking like that too.

Re: Choosing an AI model: one prompt, 11 models, different results

#103

> Build a one-page site for a neighbourhood coffee shop: opening hours, the address, a short menu and a photo. Nothing on it changes unless I edit it myself. If that's the entire prompt, it's quite depressing how much alike these all look. I appreciate some of the details from the Opus 5 version, but I can't help but strongly feel the AI vibes emanating from that design.

Reminds me of the Bootstrap era. I can't count the amount of websites with a centered header navigation, slightly rounded accent color buttons, hero section that was mostly text, and some "fun quirk" in the background, either geometric shapes or squiggles or something.

I think websites have always looked mostly alike. It's sorta always been a thing. Reminds me of "Corporate Memphis" (https://en.wikipedia.org/wiki/Corporate_Memphis)

It's alright though, because some people are OK with middle of the road (Wordpress, Boostrap, Squarespace templates, now AI). And others are willing to either pay a developer to get involved or put in the extra effort to differentiate themselves.

Re: Choosing an AI model: one prompt, 11 models, different results

#104

Earlier quoted context omitted.

When I travel I have limited time, and if e.g. the coffee shop doesn't have the hallmarks of good coffee, I will save my time and find one that does

So you compare by visiting every website of each coffee shop? And that saves you time?

Why are you trying to put words in my mouth about visiting every coffee shops website etc? You look at possible places and choose the one(s) that appeal.

Since all coffee stores are not the same, I look at the website of the coffee shop I am considering visiting, and either I choose it or I don't.

Re: Choosing an AI model: one prompt, 11 models, different results

#105
I'm doing a small local version of this test to pull moods and themes out of song lyrics. I've found that one prompt, 1 model, 11 runs even gives different results.

There's no consistency over multiple runs of the same prompt on the same model.

Also, for the purpose of music lyric analysis Qwen 4B is laughably bad. Like it's going out of its way to be extremely wrong, misunderstand the prompt. When it does correctly understand what I asked for it ALWAYS tells me that the mood is Angry. Sometimesiit just gives back all the lyrics. Sometimes it claims that it doesn't have the list of moods or the lyrics and tells me I should look them up on the Internet first.

All with the same prompt every time.

Are models being overtuned for coding tasks?

Post reply on HN