I've written about Humanity's Last Exam, which crowdsources tough questions for AI models from domain experts around the world. https://www.happiesthealth.com/articles/future-of-health/hum... It's a shifting goalpost, but one of the things that struck me was how some questions could still be trivial for a fairly qualified human (a doctor in this case) but difficult for an AI model. Reasoning, visual or logic, is buil…
Study identifies weaknesses in how AI systems are evaluated
131–140 of 204 posts
Re: Study identifies weaknesses in how AI systems are evaluated
#132Benchmarks are like SAT scores. Can they guarantee you'll be great at your future job? No, but we are still roughly okay with what they signify. Clearly LLMs are getting better in meaningful ways, and benchmarks correlate with that to some extend.
There’s no a priori reason to expect a test designed to test human academic performance would be a good one to test LLM job performance. For example a test of “multiply 1765x9392” would have some correlation with human intelligence but it wouldn’t make sense to apply it to computers.
Re: Study identifies weaknesses in how AI systems are evaluated
#133Benchmarks are like SAT scores. Can they guarantee you'll be great at your future job? No, but we are still roughly okay with what they signify. Clearly LLMs are getting better in meaningful ways, and benchmarks correlate with that to some extend.
We took objective computers, and made them generate subjective results. Isn’t this a problem that we already know there’s no solution to?
That grading subjectivity is just subjective itself.
Re: Study identifies weaknesses in how AI systems are evaluated
#134Earlier quoted context omitted.
There’s no a priori reason to expect a test designed to test human academic performance would be a good one to test LLM job performance. For example a test of “multiply 1765x9392” would have some correlation with human intelligence but it wouldn’t make sense to apply it to computers.
Actually… ask gpt1 to multiply 1765x9392.
There are LLMs, the engines that make these products run, and then the products themselves.
GPT anything should not be asked math problems. LLMs are language models, not math.
The line is going to get very blurry because ChatGPT, or Claude or Gemini, are not LLM’s. Their products driven by LLMs.
The question or requisite should not be can my LLM do math. It can I build a product that is LLM driven that can reason through math problems. Those are different things.
A coworker of mine told me that GPT’s LLM can use Excel files. No, it can’t. But the tools they plugged into it can.
Re: Study identifies weaknesses in how AI systems are evaluated
#135(Don't hurt me, I just like his chatbot. It's the best I've tried at, "Find the passage in X that reminded me of the passage in Y given this that and the other thing." It has a tendency to blow smoke if you let it, but they all seek to affirm more than I'd like, but ain't that the modern world? It can also be hilariously funny in surprisingly apt ways.)
Re: Study identifies weaknesses in how AI systems are evaluated
#136They should laugh while they can ;) Still waiting for the crash and to see what lives on and what gets recycled. My bet is that grok is here to stay ;) (Don't hurt me, I just like his chatbot. It's the best I've tried at, "Find the passage in X that reminded me of the passage in Y given this that and the other thing." It has a tendency to blow smoke if you let it, but they all seek to affirm more than I'd like, but a…
Re: Study identifies weaknesses in how AI systems are evaluated
#137They should laugh while they can ;) Still waiting for the crash and to see what lives on and what gets recycled. My bet is that grok is here to stay ;) (Don't hurt me, I just like his chatbot. It's the best I've tried at, "Find the passage in X that reminded me of the passage in Y given this that and the other thing." It has a tendency to blow smoke if you let it, but they all seek to affirm more than I'd like, but a…
Grok is terrible at coding though.
Re: Study identifies weaknesses in how AI systems are evaluated
#138Re: Study identifies weaknesses in how AI systems are evaluated
#139Benchmarks are like SAT scores. Can they guarantee you'll be great at your future job? No, but we are still roughly okay with what they signify. Clearly LLMs are getting better in meaningful ways, and benchmarks correlate with that to some extend.