OpenAI Models Dominate Structured Code Edit Benchmark
blog.mentat.ai
OpenAI Models Dominate Structured Code Edit Benchmark
1–2 of 2 posts
Re: OpenAI Models Dominate Structured Code Edit Benchmark
#2I saw the same thing with Mailogy.
I only tested variants of gpt-3.5 and -4 but got ~50% invalid syntax errors with 3.5, and virtually none with 4.