Live data from Hacker News

MAI-Code-1-Flash

microsoft.ai

11–20 of 297 posts

Re: MAI-Code-1-Flash

#17
post #6

Earlier quoted context omitted.

What is your evidence for this claim?

They say hill climbing https://microsoft.ai/news/building-a-hillclimbing-machine-la... Unless they specifically clarify that the testing and training benchmarks are completely separate, we have to assume they test on the same 'hill' the model climbs.

[flagged]

Re: MAI-Code-1-Flash

#18
is 51% good enough to reliably use? There's no world in which I use an AI agent where it gets even 15% of the code wrong, that's as bad a Tesla FSD where you need to pay attention to the road while engaging FSD. What's the point? My attention is what I'm trying to relieve, not mostly correct functionality. The only thing that matters is whether you can one-shot code like Claude or Codex, I'm not interested in a small but mostly-okay-but-annoyingly-buggy-every-now-and-then AI.
Post reply on HN