So, the model is bad at helping in this particular task. How does this compare with a control of a beneficial human task? Like someone in a lab testing blood samples or working on cancer research? Is the model equally useless for those types of lab tasks? What about other complex tasks, like home repair or architecture? Is this a success of guardrails or a failing of the model in general?
Here's what LLMs are good for: * Taking care of boilerplate work for people who know what they are doing (somewhat unreliably) * Brainstorming ideas for people who know what they are doing * Making people who don't quite know what they're doing look like they know what they're doing a little better (somewhat unreliably) LLMs are like having an army of very knowledgable but somewhat senseless interns to do your biddin…
I prefer to think of them as the underwear gnomes, just more widely read and better at BS-ing.
What happens when everyone gets to have a tireless army of very knowledgeable and AVERAGE common sense interns who have brains directly wired to various software tools, working 24/7 at 5X the speed? In the hands of a highly motivated rogue organization, this could be quite dangerous.
This is a bit beyond where we are now, but shouldn't we be prepared for this ahead of time?