Am I wrong or are these evaluations, while interesting, not really meaningful for anyone doing serious development work? I'm asking because I personally only use AI with specific and detailed instructions, building my projects piece-by-piece. I mostly don't look at the low level code and some of it I don't understand as much as I'd like, but I very much give much more technical instructions than a simple, two sentenc…
I've been using code review as my benchmark. Launched a massive parallel Fable/max code review on my lone lisp codebase and recorded all the results and transcripts. Switched to OpenAI and am now running the exact same code review with Sol/max. It's still a work in progress but preliminary results reveal that Sol is able to reproduce 70%-90% of Fable's performance. This is a very meaningful result for me because code…
Choosing an AI model: one prompt, 11 models, different results
91–100 of 105 posts
Re: Choosing an AI model: one prompt, 11 models, different results
#92Earlier quoted context omitted.
The obvious limitation to this approach is that you're limited to doing evals for things where the SOTA models are capable of judging an attempt. Many tasks are not yet amenable to this, especially if you don't have a ground truth to have the SOTA model compare to. And even for the tasks that are amenable to having LLM judges assess them, there's still a huge benefit in looking at traces yourself and labeling them. I…
Yes, that's true. There are many tasks you cannot do this on, but it's far beyond reading other people's broad benchmarks IMHO. e.g. there are tasks where Gemma is great, but on my camera watching, Qwen 3.6's MoE far exceeds either the Gemma 4 Dense or MoE. I suppose the point I meant to make is that evals for a large category of tasks are so accessible that other people's benchmarks are not useful at all.
Re: Choosing an AI model: one prompt, 11 models, different results
#93Earlier quoted context omitted.
To see the hours, get a sense of the menu. There are many matcha shops by me for instance and my wife likes to check the specials before choosing which one to go to.
Doesn't most of this show up on google anyway?
Re: Choosing an AI model: one prompt, 11 models, different results
#94Am I wrong or are these evaluations, while interesting, not really meaningful for anyone doing serious development work? I'm asking because I personally only use AI with specific and detailed instructions, building my projects piece-by-piece. I mostly don't look at the low level code and some of it I don't understand as much as I'd like, but I very much give much more technical instructions than a simple, two sentenc…
As someone who's done web development for 20+ years, I find the model personalities pretty dang fascinating, especially how they develop (and evolve) design sensibilities. I'm interested to know how the classic AI "purple preference" emerged (organically?) and if the beige wave came out of specific training efforts to combat it? To your point on development work (the code itself), I was talking to some friends on the…
Re: Choosing an AI model: one prompt, 11 models, different results
#95Re: Choosing an AI model: one prompt, 11 models, different results
#96Earlier quoted context omitted.
Where do you think Google and co. source that data from? Having an actual, searchable, up to date menu is, just for a11y and accuracy far more valuable than someone’s poorly lit, low res upload of an photo taken from the menu circa 2019, just to name one advantage. Have found data from restaurants without their own webpages on Maps utterly inaccurate, even suggesting some that have been shut down for months to where…
If you are a business, you can set up google's business feature and provide the data and high-quality images. That is basic digital presence hygeine. If a shop is refusing to do that, they will definitely not care about a website
Re: Choosing an AI model: one prompt, 11 models, different results
#97Re: Choosing an AI model: one prompt, 11 models, different results
#98Earlier quoted context omitted.
To see the hours, get a sense of the menu. There are many matcha shops by me for instance and my wife likes to check the specials before choosing which one to go to.
Doesn't most of this show up on google anyway?
Re: Choosing an AI model: one prompt, 11 models, different results
#99> Build a one-page site for a neighbourhood coffee shop: opening hours, the address, a short menu and a photo. Nothing on it changes unless I edit it myself. If that's the entire prompt, it's quite depressing how much alike these all look. I appreciate some of the details from the Opus 5 version, but I can't help but strongly feel the AI vibes emanating from that design.
I trialed Claude for a sorely needed redesign on my website. I was quite impressed with the results, but decided to hang fire on deploying it. Cue my surprise when I just found the same design applied to a coffee shop! The similarities are beyond coincidence, to the point I'll be scrapping Claude's version of the redesign.
Re: Choosing an AI model: one prompt, 11 models, different results
#100> Build a one-page site for a neighbourhood coffee shop: opening hours, the address, a short menu and a photo. Nothing on it changes unless I edit it myself. If that's the entire prompt, it's quite depressing how much alike these all look. I appreciate some of the details from the Opus 5 version, but I can't help but strongly feel the AI vibes emanating from that design.
It's also kind of pointless. Why does a coffee shop need a website anyway? Nobody is saying, "man I would love to go to this coffee shop but I can't find their website". If all you want is "opening hours, the address, a short menu and a photo" there are easier and cheaper ways to do that.
And believe it or not, pre-ground coffee has a risk of gluten cross contamination. That's not a problem with most coffee shops though as long as I get an espresso instead of drip coffee