Live data from Hacker News

Study identifies weaknesses in how AI systems are evaluated

oii.ox.ac.uk

131–140 of 204 posts

Re: Study identifies weaknesses in how AI systems are evaluated

#131

I've written about Humanity's Last Exam, which crowdsources tough questions for AI models from domain experts around the world. https://www.happiesthealth.com/articles/future-of-health/hum... It's a shifting goalpost, but one of the things that struck me was how some questions could still be trivial for a fairly qualified human (a doctor in this case) but difficult for an AI model. Reasoning, visual or logic, is buil…

Mercor is doing doing nine digit per year revenue doing just that. Micro1 and others also.

Re: Study identifies weaknesses in how AI systems are evaluated

#132

Benchmarks are like SAT scores. Can they guarantee you'll be great at your future job? No, but we are still roughly okay with what they signify. Clearly LLMs are getting better in meaningful ways, and benchmarks correlate with that to some extend.

There’s no a priori reason to expect a test designed to test human academic performance would be a good one to test LLM job performance. For example a test of “multiply 1765x9392” would have some correlation with human intelligence but it wouldn’t make sense to apply it to computers.

Actually… ask gpt1 to multiply 1765x9392.

Re: Study identifies weaknesses in how AI systems are evaluated

#133

Benchmarks are like SAT scores. Can they guarantee you'll be great at your future job? No, but we are still roughly okay with what they signify. Clearly LLMs are getting better in meaningful ways, and benchmarks correlate with that to some extend.

Isn’t this like grading art critics?

We took objective computers, and made them generate subjective results. Isn’t this a problem that we already know there’s no solution to?

That grading subjectivity is just subjective itself.

Re: Study identifies weaknesses in how AI systems are evaluated

#134

Earlier quoted context omitted.

There’s no a priori reason to expect a test designed to test human academic performance would be a good one to test LLM job performance. For example a test of “multiply 1765x9392” would have some correlation with human intelligence but it wouldn’t make sense to apply it to computers.

Actually… ask gpt1 to multiply 1765x9392.

I wish this was more broadly, explained to people…

There are LLMs, the engines that make these products run, and then the products themselves.

GPT anything should not be asked math problems. LLMs are language models, not math.

The line is going to get very blurry because ChatGPT, or Claude or Gemini, are not LLM’s. Their products driven by LLMs.

The question or requisite should not be can my LLM do math. It can I build a product that is LLM driven that can reason through math problems. Those are different things.

A coworker of mine told me that GPT’s LLM can use Excel files. No, it can’t. But the tools they plugged into it can.

Re: Study identifies weaknesses in how AI systems are evaluated

#135
They should laugh while they can ;) Still waiting for the crash and to see what lives on and what gets recycled. My bet is that grok is here to stay ;)

(Don't hurt me, I just like his chatbot. It's the best I've tried at, "Find the passage in X that reminded me of the passage in Y given this that and the other thing." It has a tendency to blow smoke if you let it, but they all seek to affirm more than I'd like, but ain't that the modern world? It can also be hilariously funny in surprisingly apt ways.)

Re: Study identifies weaknesses in how AI systems are evaluated

#136

They should laugh while they can ;) Still waiting for the crash and to see what lives on and what gets recycled. My bet is that grok is here to stay ;) (Don't hurt me, I just like his chatbot. It's the best I've tried at, "Find the passage in X that reminded me of the passage in Y given this that and the other thing." It has a tendency to blow smoke if you let it, but they all seek to affirm more than I'd like, but a…

Grok is terrible at coding though.

Re: Study identifies weaknesses in how AI systems are evaluated

#137

They should laugh while they can ;) Still waiting for the crash and to see what lives on and what gets recycled. My bet is that grok is here to stay ;) (Don't hurt me, I just like his chatbot. It's the best I've tried at, "Find the passage in X that reminded me of the passage in Y given this that and the other thing." It has a tendency to blow smoke if you let it, but they all seek to affirm more than I'd like, but a…

Grok is terrible at coding though.

[deleted]

Re: Study identifies weaknesses in how AI systems are evaluated

#139

Benchmarks are like SAT scores. Can they guarantee you'll be great at your future job? No, but we are still roughly okay with what they signify. Clearly LLMs are getting better in meaningful ways, and benchmarks correlate with that to some extend.

People often use "clearly" or "obviously" to elide the subject that is under discussion. People are saying that they do not think that it is clear that LLMs are getting better in meaningful ways, and they are saying that the benchmarks are problematic. "Clearly" isn't a counterargument.
Post reply on HN