Live data from Hacker News

Apple's On-Device and Server Foundation Models

machinelearning.apple.com

41–50 of 562 posts

Re: Apple's On-Device and Server Foundation Models

#41
post #5

Earlier quoted context omitted.

So built on stolen data essentially.

Does that imply I just stole your comment by reading it? No snark intended; I’m seriously asking. If the answer is “no” then where do you draw the line?

Data gets either stolen or freed depending on whether the guy who copied it is someone you dislike or like. Personally, I think that Apple is giving the data more exposure which, as I've been informed many times here, is much more valuable than paying for the data.

Re: Apple's On-Device and Server Foundation Models

#42

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

Apple just did more to make this a privacy focused feature versus just a data mine than literally anyone else to date and still people complain. Public content on the internet is public content on the internet - I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.

Oh no, don't get me wrong. I like the privacy features, it's already way better than OpenAI's "we make it proprietary so we can spy on you" approach.

What I don't like is the hypocrisy that basically every AI company has engaged in, where copying my shit is OK but copying theirs is not. The Internet is not public domain, as much as Eric Bauman and every AI research team would say otherwise. Even if you don't like copyright[0], you should care about copyleft, because denying valuable creative work to the proprietary world is how you get them to concede. If you can shove that work into an AI and get the benefits of that knowledge without the licensing requirement, then copyleft is useless as a tactic to get the proprietary world to bend the knee.

[0] And I don't.

My opinion is that individual copyright ownership is a bad deal for most artists and we need collective negotiation instead. Even the most copyright-respecting, 'ethical' AI boils down to Adobe dropping a EULA roofie in the Adobe Stock Contributor Agreement that lets them pay you pennies.

Re: Apple's On-Device and Server Foundation Models

#43
post #35

Earlier quoted context omitted.

> And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement where they've already pirated shittons of data This is wrong. AppleBot identifier hasn't changed: https://support.apple.com/en-us/119829 There is no AppleBot-Extended. And if you blocked it in the past it remains blocked.

From your own link: > Controlling data usage > In addition to following all robots.txt rules and directives, Apple has a secondary user agent, Applebot-Extended, that gives web publishers additional controls over how their website content can be used by Apple. > With Applebot-Extended, web publishers can choose to opt out of their website content being used to train Apple’s foundation models powering generative AI fe…

Might want to actually read it:

Applebot-Extended does not crawl webpages.

They gave this as an additional control to allow crawling for search but blocking for use in models.

Re: Apple's On-Device and Server Foundation Models

#44
post #35

Earlier quoted context omitted.

From your own link: > Controlling data usage > In addition to following all robots.txt rules and directives, Apple has a secondary user agent, Applebot-Extended, that gives web publishers additional controls over how their website content can be used by Apple. > With Applebot-Extended, web publishers can choose to opt out of their website content being used to train Apple’s foundation models powering generative AI fe…

Might want to actually read it: Applebot-Extended does not crawl webpages. They gave this as an additional control to allow crawling for search but blocking for use in models.

> There is no AppleBot-Extended. And if you blocked it in the past it remains blocked.

You said there is no Applebot-Extended. The link says otherwise.

Re: Apple's On-Device and Server Foundation Models

#45
post #22

Earlier quoted context omitted.

Does that imply I just stole your comment by reading it? No snark intended; I’m seriously asking. If the answer is “no” then where do you draw the line?

Reading, no. Selling derivative works using, yes.

If I read your comment, then write a reply, is it a derivative work?

Re: Apple's On-Device and Server Foundation Models

#46
post #7

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

> just trained a new OS development AI on every OS Apple has ever written. …is there publicly visible source code for every OS Apple has ever written?

Partially:

https://github.com/apple-oss-distributions/distribution-macO...

https://github.com/apple-oss-distributions/distribution-iOS

I'm not sure how it all fits together but people have even made an open source distribution of the base of darwin, the underlying OS:

https://github.com/PureDarwin/PureDarwin

Re: Apple's On-Device and Server Foundation Models

#47
post #35

Earlier quoted context omitted.

> And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement where they've already pirated shittons of data This is wrong. AppleBot identifier hasn't changed: https://support.apple.com/en-us/119829 There is no AppleBot-Extended. And if you blocked it in the past it remains blocked.

From your own link: > Controlling data usage > In addition to following all robots.txt rules and directives, Apple has a secondary user agent, Applebot-Extended, that gives web publishers additional controls over how their website content can be used by Apple. > With Applebot-Extended, web publishers can choose to opt out of their website content being used to train Apple’s foundation models powering generative AI fe…

But it also says that Applebot-Extended doesn't crawl webpages and instead this marker is only used to determine what can be done with the pages that were visited by Applebot.

Not that I like an opt-out system, but based on the wording of the docs it is true that if you blocked Applebot then blocking Applebot-Extended isn't necessary.

Re: Apple's On-Device and Server Foundation Models

#48

Earlier quoted context omitted.

Likely they’ll be able to take advantage of the hardware neural engine and be far more power efficient. Apple has demonstrated this is something it takes pretty seriously.

So iOS LLM Apps dont use the neural engine? Lol

Probably not. The CoreML LLM stuff only works on Macs AFAIK. Probably the phone app uses the GPU.

Re: Apple's On-Device and Server Foundation Models

#49
post #5

Earlier quoted context omitted.

So built on stolen data essentially.

Web scraping is legal. And if you run a website and want to opt-out then simply add a robots.txt. The standard way of preventing bots for 30 years.

How are people supposed to block it when they stole all the data first and then only after that point they decide to even tell anyone what user agent they need to block and how they are planning to exploit your work for their profit.

Re: Apple's On-Device and Server Foundation Models

#50
post #20

Earlier quoted context omitted.

Does that imply I just stole your comment by reading it? No snark intended; I’m seriously asking. If the answer is “no” then where do you draw the line?

I don’t actually think this is complicated and reading a comment is not the same thing as scraping the internet and you obviously know that. A few factors that come to mind would be: - scale - informed consent which there was none in this case - how you are going to use that data. For example using everybody others work so the worlds richest company can make more money from it while giving back nothing in return is a…

Reading a comment is exactly the same thing as scraping the internet, you just stop sooner.
Post reply on HN