Live data from Hacker News

DeepSeek may have used Google's Gemini to train its latest model

techcrunch.com

1–10 of 17 posts

Re: DeepSeek may have used Google's Gemini to train its latest model

#3
post #2

At this point, they all using each other because so much of the new content they are scraping for data is generated. These models will converge and plateau because the datasets are only going to get worse as more of their content is incestuous.

Yes indeed some studies were already done on this.

Re: DeepSeek may have used Google's Gemini to train its latest model

#4
> Distillation isn’t an uncommon practice, but OpenAI’s terms of service prohibit customers from using the company’s model outputs to build competing AI.

I have the absolute tiniest of violins for this given OpenAI's behaviour vs everyone else's terms of service.

Re: DeepSeek may have used Google's Gemini to train its latest model

#5

> Distillation isn’t an uncommon practice, but OpenAI’s terms of service prohibit customers from using the company’s model outputs to build competing AI. I have the absolute tiniest of violins for this given OpenAI's behaviour vs everyone else's terms of service.

“Copyright must evolve into the 21century (…so that AI can legally steal everything produced by people”

And also “Don’t steal our AI!”

Re: DeepSeek may have used Google's Gemini to train its latest model

#7
post #2

At this point, they all using each other because so much of the new content they are scraping for data is generated. These models will converge and plateau because the datasets are only going to get worse as more of their content is incestuous.

I recall that AI trained on AI output over many cycles eventually becomes something akin to noise texture as the output degrades rapidly.

Won’t most AI produced content put out into the public be human curated, thus heavily mitigating this degradation effect? If we’re going to see a full length AI generated movie it seems like humans will be heavily involved, hand holding the output and throwing out the AI’s nonsense.

Re: DeepSeek may have used Google's Gemini to train its latest model

#8

> Distillation isn’t an uncommon practice, but OpenAI’s terms of service prohibit customers from using the company’s model outputs to build competing AI. I have the absolute tiniest of violins for this given OpenAI's behaviour vs everyone else's terms of service.

I'm still unclear how they are able to claim this considering their raw thinking traces were never exposed to the end user, only summaries.

Re: DeepSeek may have used Google's Gemini to train its latest model

#9
post #2

At this point, they all using each other because so much of the new content they are scraping for data is generated. These models will converge and plateau because the datasets are only going to get worse as more of their content is incestuous.

The default Llama 4 system prompt even instructs it to avoid using various ChatGPT-isms, presumably because they've already scraped so much GPT-generated material that it noticably skews their models output.

Re: DeepSeek may have used Google's Gemini to train its latest model

#10

> Distillation isn’t an uncommon practice, but OpenAI’s terms of service prohibit customers from using the company’s model outputs to build competing AI. I have the absolute tiniest of violins for this given OpenAI's behaviour vs everyone else's terms of service.

“Copyright must evolve into the 21century (…so that AI can legally steal everything produced by people” And also “Don’t steal our AI!”

The world is not prepared for the mental gymnastics that OpenAI/Google/etc will employ to defend their copyright if their big models ever get leaked.
Post reply on HN