Live data from Hacker News

Study identifies weaknesses in how AI systems are evaluated

oii.ox.ac.uk

201–204 of 204 posts

Re: Study identifies weaknesses in how AI systems are evaluated

#201

Earlier quoted context omitted.

Sorry, pal, but if benchmarks were to disagree with opinions of a bunch of users saying "tech companies bad"? I'd side with benchmarks at least 9 times out of 10.

How does that have anything to do with what we're talking about?

What that has to do is: your "tech companies are bad for using literally the best tool we have for measuring AI capabilities when talking about AI capablities" take is a very bad take.

It's like you wanted to say "tech companies are bad", and the rest is just window dressing.

Re: Study identifies weaknesses in how AI systems are evaluated

#202
On top of this, we have complex infrastructure setups with different TPUs and GPUs across multiple data centers. The benchmarks tested when models were released might not reflect what we're actually using now. We need to evaluate models continuously, for example, https://isitnerfed.org/ does exactly that.

Re: Study identifies weaknesses in how AI systems are evaluated

#203
post #49

Earlier quoted context omitted.

we try to make benchmarks for users, but it's like that 20% article - different people want different 20% and you just end up adding "features" and whackamoling the different kinds of 20% if a single benchmark could be a universal truth, and it was easy to figure out how to do it, everyone would love that.. but that's why we're in the state we're in right now

The problem isn’t with the benchmarks (or the models, for that matter) it’s their being used to prop up the indefensible product marketing claims made by people frantically justifying asking for more dump trucks of thousand-dollar bills to replace the ones they just burned through in a few months.

unfortunately as benchmark makers we can't really do anything about human nature :shrug:

Re: Study identifies weaknesses in how AI systems are evaluated

#204

Earlier quoted context omitted.

The problem isn’t with the benchmarks (or the models, for that matter) it’s their being used to prop up the indefensible product marketing claims made by people frantically justifying asking for more dump trucks of thousand-dollar bills to replace the ones they just burned through in a few months.

unfortunately as benchmark makers we can't really do anything about human nature :shrug:

Absolutely not. This is not a problem with any part of the engineering process. Nearly everything wrong with the AI business lies at the feet of product managers, marketing, the c-suite crowd, etc.
Post reply on HN