Live data from Hacker News

MAI-Thinking-1

microsoft.ai

11–20 of 90 posts

Re: MAI-Thinking-1

#12
post #7

Earlier quoted context omitted.

I'm interested how much "Clean Data" is synthetic data from "unclean" models...

“ We trained it from the ground up on enterprise grade, clean and commercially licensed data, without distillation from third-party models.”

aka all of GitHub OSS

Re: MAI-Thinking-1

#13
post #5

> Second, clean data. MAI-Thinking-1 was trained on clean and appropriately licensed data, with AI-generated content excluded from pre-training. This matters for quality, provenance, and control. If we cannot account for what shaped a model, we cannot fully understand its behavior or credibly improve it. Shots fired? It would be interesting to see how far "clean data" can go on the scaling laws.

I'm interested how much "Clean Data" is synthetic data from "unclean" models...

> with AI-generated content excluded from pre-training.

> without distillation from third-party models

sounds like zero unless they are lying.

Re: MAI-Thinking-1

#14
They've hijacked scrolling. They've hijacked the spacebar. It flickers like crazy when I try to move through the article. Trying to get through it is an exercise in madness.

Re: MAI-Thinking-1

#15
post #5

> Second, clean data. MAI-Thinking-1 was trained on clean and appropriately licensed data, with AI-generated content excluded from pre-training. This matters for quality, provenance, and control. If we cannot account for what shaped a model, we cannot fully understand its behavior or credibly improve it. Shots fired? It would be interesting to see how far "clean data" can go on the scaling laws.

I doubt any lab would say otherwise, they all _claim_ to use licensed data

Re: MAI-Thinking-1

#17

They've hijacked scrolling. They've hijacked the spacebar. It flickers like crazy when I try to move through the article. Trying to get through it is an exercise in madness.

I normally don't comment on matters of taste like this, but wow this is brutal. It's like someone threw the site in a vat of molasses.

Re: MAI-Thinking-1

#18
post #13

Earlier quoted context omitted.

I'm interested how much "Clean Data" is synthetic data from "unclean" models...

> with AI-generated content excluded from pre-training. > without distillation from third-party models sounds like zero unless they are lying.

> with AI-generated content excluded from pre-training.

Though this is largely impossible these days, unless they pre-trained on pre-AI era data.

Re: MAI-Thinking-1

#19

Looks like the OAI divergence is finally taking place. Seems like the comparisons are mainly with Opus 4.6 and GPT 5.4 though. Still, exciting to see a new frontier player.

Is it a frontier player though, or perhaps a new benchmaxxed model? People were saying similar things about Grok but it ultimately amounted to little.

Re: MAI-Thinking-1

#20

They've hijacked scrolling. They've hijacked the spacebar. It flickers like crazy when I try to move through the article. Trying to get through it is an exercise in madness.

I do not understand how scroll hijacking is still a thing. Who thinks this is a better experience?
Post reply on HN