Earlier quoted context omitted.
I think this is still useful research that calls into question how “smart” these models are. If the model needs a separate tool to solve a problem, has the model really solved the problem, or just outsourced it to a harness that it’s been trained - via reinforcement learning - to call upon?
Does it matter if the LLM can solve the problem or if it knows to use a resource? There’s plenty of math that I couldn’t even begin to solve without a calculator or other tool. Doesn’t mean I’m not solving math problems. In woodworking, the advice is to let the tool do the work. Does someone using a power saw have less claim to having built something than a handsaw user? Does a CNC user not count as a woodworker beca…
Re: The case for zero-error horizons in trustworthy LLMs
#121Is your issue with math in this example the tediousness of the operations or a conceptual lack of understanding of how to solve them?