Live data from Hacker News

Apple's On-Device and Server Foundation Models

machinelearning.apple.com

101–110 of 562 posts

Re: Apple's On-Device and Server Foundation Models

#102

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

> And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement where they've already pirated shittons of data.

It’s not as bad as that, I think. https://support.apple.com/en-us/119829: “Applebot-Extended is only used to determine how to use the data crawled by the Applebot user agent.“

⇒ if you use robots.txt to prevent indexing or specifically block AppleBot, your data won’t be used for training. AppleBot is almost a decade old (https://searchengineland.com/apple-confirms-their-web-crawle...)

Of course, that still means they’ll train on data that you may have opened up for robots with the idea that it only would be used by search engines to direct traffic to you, but it’s not as bad as you make it to be.

Re: Apple's On-Device and Server Foundation Models

#105

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

[deleted]

Re: Apple's On-Device and Server Foundation Models

#106
post #26

Earlier quoted context omitted.

> publicly available data collected Data, implies factual information. You can not copyright factual information. The fact that I use the word "appalling" to describe the practice of doing this results in some vector relationship between the words. Thats the data, the fact, not the writing itself. There are going to be a bunch of interesting court cases where the court is going to have to backtrack on copyrighting fa…

> Data, implies factual information. You can not copyright factual information Where on Earth did you get that from?

> "data implies factual information"

They used the word DATA, not content, DATA...

The argument that is going to be made, that your copy right work stands. That the model doesn't care about your document it cares that "the" was used N number of times and its relationships to other words. That information isnt your work, and it is factual. That "data" only has value is when it's weighted against all the "data" put into the system, again not your work at all. (We would say thats information derived, but it will be argued that it is transformed).

> You can not copyright factual information

https://www.techdirt.com/2007/11/27/yet-again-court-tells-ml...

The MLB has been trying to copyright baseball stats forever. The court keeps saying "you cant copyright facts".

Re: Apple's On-Device and Server Foundation Models

#107

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

Apple just did more to make this a privacy focused feature versus just a data mine than literally anyone else to date and still people complain. Public content on the internet is public content on the internet - I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.

You seen to misunderstand what licensing , or ‘unlicensed’, actually means.

If I write a story a publish it freely on line to my website it’s not ’unlicensed’ in a way that means anyone had the right to yank it and republish it. Even though it’s freely available, I still own the copyright of it.

Similarly, we don’t say that GPL-ed code is ‘unlicensed’ just because it is available for free. It has a license, which defines very specific terms that must be followed.

Re: Apple's On-Device and Server Foundation Models

#108
post #81
post #71

Earlier quoted context omitted.

Where did you get the idea that's its way better than openai's? Aren't they both proprietary?

Apple isn’t collecting data from their customers. Edit: to feed back into their AI training.

Apple necessarily collects data from customers for mandatory reasons, like everyone else. (Like, you need someone's address to ship them their order.)

More useful questions are if they're using it for other purposes without opt-in or accidentally leaking it.

Re: Apple's On-Device and Server Foundation Models

#109

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

Apple just did more to make this a privacy focused feature versus just a data mine than literally anyone else to date and still people complain. Public content on the internet is public content on the internet - I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.

I’m sorry but I really dislike this perspective. “Every one else has been awful. Apple is being less awful and you’re still complaining?”

Yeah, I’m complaining. We all agreed years ago to web indexing conventions still in practise today. No, no one is obliged to follow them but you can rest assured I’ll complain about them. There was a time when the web felt like a cooperative place, these days it’s just value extraction after value extraction.

Re: Apple's On-Device and Server Foundation Models

#110

Halfway down the article contains some great charts with comparisons to other relevant models, like Mistral-7B for the on-device models, and both gpt-3.5 and 4 for the server-side models. They include data about the ratio of which outputs human graders preferred (for server side it’s better than 3.5, worse than 4). BUT, the interesting chart to me is „Human Evaluation of Output Harmfulness” which is much, much ”bette…

I want to know what they consider "harmful". Is it going to refuse to operate for sex workers, murder mystery writers, or people who use knives?

The caption for the image gives a little more insight into "harmful" and one of the things it mentions is factuality - which is interesting, but doesn't reveal a whole lot unless they were to break it out by "type of harmful".
Post reply on HN