Earlier quoted context omitted.
> with AI-generated content excluded from pre-training. > without distillation from third-party models sounds like zero unless they are lying.
> with AI-generated content excluded from pre-training. Though this is largely impossible these days, unless they pre-trained on pre-AI era data.
MAI-Thinking-1
71–80 of 90 posts
Re: MAI-Thinking-1
#72Re: MAI-Thinking-1
#73Re: MAI-Thinking-1
#74> Second, clean data. MAI-Thinking-1 was trained on clean and appropriately licensed data, with AI-generated content excluded from pre-training. This matters for quality, provenance, and control. If we cannot account for what shaped a model, we cannot fully understand its behavior or credibly improve it. Shots fired? It would be interesting to see how far "clean data" can go on the scaling laws.
I would really like to see what "appropriately licensed data" means. Cannot imagine they didn't copy all open repo's on GitHub, and can't imagine they asked for permission, or are reproducing license texts from these repo's now. It sounds hand wavy. P.S. A fairly basic website otherwise, but it unfortunately seems to be hacking scroll for no good reason.
Re: MAI-Thinking-1
#75Earlier quoted context omitted.
I would really like to see what "appropriately licensed data" means. Cannot imagine they didn't copy all open repo's on GitHub, and can't imagine they asked for permission, or are reproducing license texts from these repo's now. It sounds hand wavy. P.S. A fairly basic website otherwise, but it unfortunately seems to be hacking scroll for no good reason.
Recently, GitHub has changed their terms of service to use all user data for AI training unless users explicitly opt out. This is probably the way Microsoft has obtained "appropriately licensed data".
Re: MAI-Thinking-1
#76I was most excited about the "frontier tuning." Like, it will actually watch you do stuff and learn to do it for you? That would be actually interesting.
But no, it's just a data labelling interface: https://learn.microsoft.com/en-us/microsoft-365/copilot/copi.... You have to provide the instruction and give feedback and there is a whole UI with hour-lonf wait between steps. So basically they want you to do the labelling to train a model, or at least that's how it looks from the outside
Also the mission statement of Humanist AI is the most boring, but tries to sound way too grand. Like "all the cool labs have a mission statement, so we should also have one" vibes
Re: MAI-Thinking-1
#77Re: MAI-Thinking-1
#78The benchmarks are a bit of a disaster? It's at about DeepSeek V3.2 level, but with about 50% more parameters. Loses handily to the also smaller GLM-5.1, and even worse to the similarly sized Kimi K2.6.
Re: MAI-Thinking-1
#79Re: MAI-Thinking-1
#80> Second, clean data. MAI-Thinking-1 was trained on clean and appropriately licensed data, with AI-generated content excluded from pre-training. This matters for quality, provenance, and control. If we cannot account for what shaped a model, we cannot fully understand its behavior or credibly improve it. Shots fired? It would be interesting to see how far "clean data" can go on the scaling laws.
I would really like to see what "appropriately licensed data" means. Cannot imagine they didn't copy all open repo's on GitHub, and can't imagine they asked for permission, or are reproducing license texts from these repo's now. It sounds hand wavy. P.S. A fairly basic website otherwise, but it unfortunately seems to be hacking scroll for no good reason.