It took me one question to have it spit out a completely dreamt up codebase, complete with emojis, promises of solutions and fixing all my problems, and of course nothing of it worked. It was a very simple question about something very well documented (Oban timeouts). I doubt LLM benchmarks more and more, what are they even testing?
It's some 3rd party thing for Elixir, a niche within a niche. I wouldn't expect an LLM to do well there.
> I doubt LLM benchmarks more and more, what are they even testing?
Probably testing by asking it to solve a problem with python or (java|type)script. Perhaps not even specifying a language and watching it generate a generic React application.