Um, what is RL?
Please, let it be Rocket Launcher for once.
11–20 of 38 posts
Um, what is RL?
Please, let it be Rocket Launcher for once.
Meta request to authors: please define your acronyms at least once! Even in scientific domains where a high level of background knowledge is expected, it is standard practice to define each acronym prior to its use in the rest of the paper, for example “using three-letter acronyms (TLAs) without first defining them is a hindrance to readability.”
> AI has beat world champions at chess and Go, surpassed most humans on SAT and bar exams, and reached gold medal level on IOI and IMO. But the world hasn’t changed much, at least judged by economics and GDP. > I call this the utility problem, and deem it the most important problem for AI. > Perhaps we will solve the utility problem pretty soon, perhaps not. Either way, the root cause of this problem might be decepti…
Earlier quoted context omitted.
I think of some of the ways LLMs perform better in real life than they do in evals. For instance I ask AI assistants a lot about what some code is trying to do in applications software where it is a matter of React, CSS and how APIs get used. Frequently this is a matter of pattern matching and doesn't require deep thought and I find LLMs often nail it. When it comes to "what does some systems oriented code do" now yo…
I think what you're describing is, easy tasks are easy to perform. Which is, of course, true. Anecdotally, a lot of value I get from Copilot is in simple, mundane tasks.
> AI has beat world champions at chess and Go, surpassed most humans on SAT and bar exams, and reached gold medal level on IOI and IMO. But the world hasn’t changed much, at least judged by economics and GDP. > I call this the utility problem, and deem it the most important problem for AI. > Perhaps we will solve the utility problem pretty soon, perhaps not. Either way, the root cause of this problem might be decepti…
To say we are at a point where AI can do anything reliably is laughable, it can do much and it will tell you any answer whether right or wrong with full confidence. To trust such a technology in the big no-human decisions like we want it to is foolswork.
Which is great! There's room in the world for new benchmarks that test for more diverse things!
It's highly likely at least one of the new benchmarks will eventually test for all the criteria being mentioned.
Meta request to authors: please define your acronyms at least once! Even in scientific domains where a high level of background knowledge is expected, it is standard practice to define each acronym prior to its use in the rest of the paper, for example “using three-letter acronyms (TLAs) without first defining them is a hindrance to readability.”