Live data from Hacker News

Study identifies weaknesses in how AI systems are evaluated

oii.ox.ac.uk

171–180 of 204 posts

Re: Study identifies weaknesses in how AI systems are evaluated

#171
post #170

Earlier quoted context omitted.

In my opinion a major weakness in how people reason about this issue is that they describe solving problems as EITHER recall of an existing solution OR creative problem solving. Sure it is possible for a specific solution to be recalled, but it's not possible for a problem to be absolutely unrelated to anything the system has ever seen before and still be solvable. There are many shades of gray in the similarity a pr…

The difference is that humans don't memorize petabytes of problems, so from a relative perspective people are constantly solving novel problems they never saw before. I'm thinking this is a requirement for dynamic, few-shot learning. We can clearly see LLMs fail when you throw even a small wrench in the prompt.

Humans encounter massive numbers of problems in their experience that informs their problem solving. The same is true of LLM. LLMs do not actually have all their training data memorized.

I’m not sure what your basis is for saying “LLMs fail if there is a small wrench in the prompt.” They also succeed despite wrenches in the prompt with great regularity.

Re: Study identifies weaknesses in how AI systems are evaluated

#172
post #170

Earlier quoted context omitted.

The difference is that humans don't memorize petabytes of problems, so from a relative perspective people are constantly solving novel problems they never saw before. I'm thinking this is a requirement for dynamic, few-shot learning. We can clearly see LLMs fail when you throw even a small wrench in the prompt.

Humans encounter massive numbers of problems in their experience that informs their problem solving. The same is true of LLM. LLMs do not actually have all their training data memorized. I’m not sure what your basis is for saying “LLMs fail if there is a small wrench in the prompt.” They also succeed despite wrenches in the prompt with great regularity.

Let's clarify, this isn't about whether the models are capable. They are very capable and impressive. This is more about whether we can use the same type of metric we use for humans to compare and conclude if they are "intelligent".

It's not just semantics, the metrics are supposed to tell us the potential of the model. If they can solve extremely hard PhD problems, it should be the case that we're already in the singularity, and they should be solving absolutely everything in whatever field they were trained in, because it's not just PhD level, it's a machine that has a ton of memory, compute and never sleeps. However, once you use these models extensively, it becomes apparent they are just synthesizing data, and not as much understanding it in a way that would allow them to extrapolate into anything else as humans do.

I think this point is a little hard to explain. I'll just emphasize, these are smart systems, and they can do a lot, but there is still a disconnect between, let's say, a PhD level model and a human with a PhD, in the "quality" of what we would call "intelligence" of both entities (human and machine).

Re: Study identifies weaknesses in how AI systems are evaluated

#173
post #153

I work on LLM benchmarks and human evals for a living in a research lab (as opposed to product). I can say: it’s pretty much the Wild West and a total disaster. No one really has a good solution, and researchers are also in a huge rush and don’t want to end up making their whole job benchmarking. Even if you could, and even if you have the right background you can do benchmarks full time and they still would be a mes…

I also work in LLM evaluation. My cynical take is that nobody is really using LLMs for stuff, and so benchmarks are mostly just make up tasks (coding is probably the exception). If we had real specific use cases it should be easier to benchmark and know if one is better, but it’s mostly all hypothetical. The more generous take is that you can’t benchmarks advanced intelligence very well, whether LLM or person. We don…

We have 20+ services in prod that use llms. So I have 50k (or more) per service per day of data to evaluate. The question is- do people actually evaluate properly.

And how do you do an apples to apples evaluation of such squishy services?

Re: Study identifies weaknesses in how AI systems are evaluated

#174
post #153

I work on LLM benchmarks and human evals for a living in a research lab (as opposed to product). I can say: it’s pretty much the Wild West and a total disaster. No one really has a good solution, and researchers are also in a huge rush and don’t want to end up making their whole job benchmarking. Even if you could, and even if you have the right background you can do benchmarks full time and they still would be a mes…

I also work in LLM evaluation. My cynical take is that nobody is really using LLMs for stuff, and so benchmarks are mostly just make up tasks (coding is probably the exception). If we had real specific use cases it should be easier to benchmark and know if one is better, but it’s mostly all hypothetical. The more generous take is that you can’t benchmarks advanced intelligence very well, whether LLM or person. We don…

You could have the world expert debate the thing. Someone who can be accused of knowing things. We have many such humans, at least as many as topics.

Publish the debate as~is so that others vaguely familiar with the topic can also be in awe or disgusted.

We have many gradients of emotion. No need to try quantify them. Just repeat the exercise.

Re: Study identifies weaknesses in how AI systems are evaluated

#175

I work on LLM benchmarks and human evals for a living in a research lab (as opposed to product). I can say: it’s pretty much the Wild West and a total disaster. No one really has a good solution, and researchers are also in a huge rush and don’t want to end up making their whole job benchmarking. Even if you could, and even if you have the right background you can do benchmarks full time and they still would be a mes…

> Brittle performance – A model might do well on short, primary school-style maths questions, but if you change the numbers or wording slightly, it suddenly fails. This shows it may be memorising patterns rather than truly understanding the problem

This finding really shocked me

Re: Study identifies weaknesses in how AI systems are evaluated

#176
post #172

Earlier quoted context omitted.

Humans encounter massive numbers of problems in their experience that informs their problem solving. The same is true of LLM. LLMs do not actually have all their training data memorized. I’m not sure what your basis is for saying “LLMs fail if there is a small wrench in the prompt.” They also succeed despite wrenches in the prompt with great regularity.

Let's clarify, this isn't about whether the models are capable. They are very capable and impressive. This is more about whether we can use the same type of metric we use for humans to compare and conclude if they are "intelligent". It's not just semantics, the metrics are supposed to tell us the potential of the model. If they can solve extremely hard PhD problems, it should be the case that we're already in the sin…

Human metrics of intelligence have always felt like rubbish. We never did this well. I would describe intelligence as effective adaption leading to survival and growth or prospering. Memorization, comprehension, speed of response etc. those are magnifying factors that are valued, we view them as components of intelligence, but llms are proving this is not the whole, without effective application, they are not intelligence. Perhaps learning is the difference? How to measure that?

Someone describing string theory is the literary equivalent of fractal structures in snowflakes. Lovely, complex, possibly unique, but not proof of a level of intelligence- for the string theorist maybe it is intelligent, perhaps persuading someone to fund their grant, which enables them to eat, shelter etc. Might be a bit harsh on string theory. Saying it is proof of an amount of intelligence leads us to falsifiable statements.

Re: Study identifies weaknesses in how AI systems are evaluated

#177

Earlier quoted context omitted.

For what it's worth, I work on platforms infra at a hyperscaler and benchmarks are a complete fucking joke in my field too lol. Ultimately we are measuring extremely measurable things that have an objective ground truth. And yet: - we completely fail at statistics (the MAJORITY of analysis is literally just "here's the delta in the mean of these two samples". If I ever do see people gesturing at actual proper analysi…

> we completely fail at statistics (the MAJORITY of analysis is literally just "here's the delta in the mean of these two samples". If I ever do see people gesturing at actual proper analysis, if prompted they'll always admit "yeah, well, we do come up with a p-value or a confidence interval, but we're pretty sure the way we calculate it is bullshit") Sort of tangential, but as someone currently taking an intro stati…

FWIW, I don't think intro stats is easy the way I normally see it taught. It focuses on formulae, tests, and step-by-step recipes without spending the time to properly develop intuition as to why those work, how they work, which ones you should use in unfamiliar scenarios, how you might find the right thing to do in unfamiliar scenarios, etc.

Pair that with skipping all the important problems (what is randomness, how do you formulate the right questions, how do you set up an experiment capable of collecting data which can actually answer those questions, etc), and it's a recipe for disaster.

It's just an exercise in box-ticking, and some students get lucky with an exceptional teacher, and others are independently able to develop the right instincts when they enter the class with the right background, but it's a disservice to almost everyone else.

Re: Study identifies weaknesses in how AI systems are evaluated

#178

I work on LLM benchmarks and human evals for a living in a research lab (as opposed to product). I can say: it’s pretty much the Wild West and a total disaster. No one really has a good solution, and researchers are also in a huge rush and don’t want to end up making their whole job benchmarking. Even if you could, and even if you have the right background you can do benchmarks full time and they still would be a mes…

The big problem is that tech companies and journalist aren't transparent about this. They tout benchmark numbers constantly, like they're an object measure of capabilities.

That's because they are as close to "object measure capabilities" as anything we're ever going to get.

Without benchmarks, you're down to evaluating model performance based on vibes and vibes only, which plain sucks. With benchmarks, you have numbers that correlate to capabilities somewhat.

Re: Study identifies weaknesses in how AI systems are evaluated

#179

Earlier quoted context omitted.

The big problem is that tech companies and journalist aren't transparent about this. They tout benchmark numbers constantly, like they're an object measure of capabilities.

That's because they are as close to "object measure capabilities" as anything we're ever going to get. Without benchmarks, you're down to evaluating model performance based on vibes and vibes only, which plain sucks. With benchmarks, you have numbers that correlate to capabilities somewhat.

That's assuming these benchmarks are the best we're ever going to get, which they clearly aren't. There's a lot to improve even without radical changes to how things are done.

Re: Study identifies weaknesses in how AI systems are evaluated

#180

Earlier quoted context omitted.

For what it's worth, I work on platforms infra at a hyperscaler and benchmarks are a complete fucking joke in my field too lol. Ultimately we are measuring extremely measurable things that have an objective ground truth. And yet: - we completely fail at statistics (the MAJORITY of analysis is literally just "here's the delta in the mean of these two samples". If I ever do see people gesturing at actual proper analysi…

“Here’s the throughout at sustained 100% load with the same ten sample queries repeated over and over.” “The customers want lower latency at 30% load for unique queries.” “Err… we can scale up for more throughput!” ಠ_ಠ

And then when you ask if they disabled the query result cache before running their benchmarking, they blink and look confused.
Post reply on HN