I love that they use machinelearning.apple.com not ai.apple.com
Apple's On-Device and Server Foundation Models
101–110 of 562 posts
Re: Apple's On-Device and Server Foundation Models
#102> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…
It’s not as bad as that, I think. https://support.apple.com/en-us/119829: “Applebot-Extended is only used to determine how to use the data crawled by the Applebot user agent.“
⇒ if you use robots.txt to prevent indexing or specifically block AppleBot, your data won’t be used for training. AppleBot is almost a decade old (https://searchengineland.com/apple-confirms-their-web-crawle...)
Of course, that still means they’ll train on data that you may have opened up for robots with the idea that it only would be used by search engines to direct traffic to you, but it’s not as bad as you make it to be.
Re: Apple's On-Device and Server Foundation Models
#103I love that they use machinelearning.apple.com not ai.apple.com
Re: Apple's On-Device and Server Foundation Models
#104Will these smaller on device models lead to a crash in GPU prices?
Re: Apple's On-Device and Server Foundation Models
#105> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…
Re: Apple's On-Device and Server Foundation Models
#106Earlier quoted context omitted.
> publicly available data collected Data, implies factual information. You can not copyright factual information. The fact that I use the word "appalling" to describe the practice of doing this results in some vector relationship between the words. Thats the data, the fact, not the writing itself. There are going to be a bunch of interesting court cases where the court is going to have to backtrack on copyrighting fa…
> Data, implies factual information. You can not copyright factual information Where on Earth did you get that from?
They used the word DATA, not content, DATA...
The argument that is going to be made, that your copy right work stands. That the model doesn't care about your document it cares that "the" was used N number of times and its relationships to other words. That information isnt your work, and it is factual. That "data" only has value is when it's weighted against all the "data" put into the system, again not your work at all. (We would say thats information derived, but it will be argued that it is transformed).
> You can not copyright factual information
https://www.techdirt.com/2007/11/27/yet-again-court-tells-ml...
The MLB has been trying to copyright baseball stats forever. The court keeps saying "you cant copyright facts".
Re: Apple's On-Device and Server Foundation Models
#107> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…
Apple just did more to make this a privacy focused feature versus just a data mine than literally anyone else to date and still people complain. Public content on the internet is public content on the internet - I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.
If I write a story a publish it freely on line to my website it’s not ’unlicensed’ in a way that means anyone had the right to yank it and republish it. Even though it’s freely available, I still own the copyright of it.
Similarly, we don’t say that GPL-ed code is ‘unlicensed’ just because it is available for free. It has a license, which defines very specific terms that must be followed.
Re: Apple's On-Device and Server Foundation Models
#108Earlier quoted context omitted.
Where did you get the idea that's its way better than openai's? Aren't they both proprietary?
Apple isn’t collecting data from their customers. Edit: to feed back into their AI training.
More useful questions are if they're using it for other purposes without opt-in or accidentally leaking it.
Re: Apple's On-Device and Server Foundation Models
#109> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…
Apple just did more to make this a privacy focused feature versus just a data mine than literally anyone else to date and still people complain. Public content on the internet is public content on the internet - I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.
Yeah, I’m complaining. We all agreed years ago to web indexing conventions still in practise today. No, no one is obliged to follow them but you can rest assured I’ll complain about them. There was a time when the web felt like a cooperative place, these days it’s just value extraction after value extraction.
Re: Apple's On-Device and Server Foundation Models
#110Halfway down the article contains some great charts with comparisons to other relevant models, like Mistral-7B for the on-device models, and both gpt-3.5 and 4 for the server-side models. They include data about the ratio of which outputs human graders preferred (for server side it’s better than 3.5, worse than 4). BUT, the interesting chart to me is „Human Evaluation of Output Harmfulness” which is much, much ”bette…
I want to know what they consider "harmful". Is it going to refuse to operate for sex workers, murder mystery writers, or people who use knives?