Live data from Hacker News

Apple's On-Device and Server Foundation Models

machinelearning.apple.com

81–90 of 562 posts

Re: Apple's On-Device and Server Foundation Models

#81
post #71

Earlier quoted context omitted.

Oh no, don't get me wrong. I like the privacy features, it's already way better than OpenAI's "we make it proprietary so we can spy on you" approach. What I don't like is the hypocrisy that basically every AI company has engaged in, where copying my shit is OK but copying theirs is not. The Internet is not public domain, as much as Eric Bauman and every AI research team would say otherwise. Even if you don't like cop…

Where did you get the idea that's its way better than openai's? Aren't they both proprietary?

Apple isn’t collecting data from their customers.

Edit: to feed back into their AI training.

Re: Apple's On-Device and Server Foundation Models

#82

Halfway down the article contains some great charts with comparisons to other relevant models, like Mistral-7B for the on-device models, and both gpt-3.5 and 4 for the server-side models. They include data about the ratio of which outputs human graders preferred (for server side it’s better than 3.5, worse than 4). BUT, the interesting chart to me is „Human Evaluation of Output Harmfulness” which is much, much ”bette…

I want to know what they consider "harmful". Is it going to refuse to operate for sex workers, murder mystery writers, or people who use knives?

Re: Apple's On-Device and Server Foundation Models

#83

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

Apple just did more to make this a privacy focused feature versus just a data mine than literally anyone else to date and still people complain. Public content on the internet is public content on the internet - I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.

Who are "we" here? Did you abolish the Berne Convention somehow?

Re: Apple's On-Device and Server Foundation Models

#85
post #20

Earlier quoted context omitted.

I don’t actually think this is complicated and reading a comment is not the same thing as scraping the internet and you obviously know that. A few factors that come to mind would be: - scale - informed consent which there was none in this case - how you are going to use that data. For example using everybody others work so the worlds richest company can make more money from it while giving back nothing in return is a…

I think it's even simpler than that: incentives. The entire premise of copyright law (and all IP law) is to protect the incentive to create new stuff, which is often a very risky and highly time or capital intensive endeavor. So here's the question: Does a person reading a comment destroy the incentive for the author to post it? No. In fact, it is the only thing that produces the incentive for someone to post. People…

> Does a model sucking up all the artistic output of the last 400 years and using that to produce an image generator model destroy the incentive of producing and sharing said artistic output?

Copyright, at least in the US, cares about the effect of the use on the market for that specific work. It's individual ownership, not collective. And while model regurgitation happens, it's less common than you think.

The real harm of AI to artists is market replacement. That is, with everyone using image generators to pop out images like candy, human artists don't have a market to sell into. This isn't even just a matter of "oh boo hoo I can't compete with Mr. Diffusion". Generative AI is very good at creating spam, which has turned every art market and social media platform into a bunch of warring spambots whose output is statistically indistinguishable from human.

The problem is, no IP law in the world is going to recognize this as a problem, because IP is a fundamentally capitalist concept. Asserting that the market for new artistic works and notoriety for those works should be the collective property of artists and artists alone is not a workable legal proposal, even if it's a valid moral principle. And conversely the history of copyright has seen it be completely subverted to the point where it only serves the interests of the publishers in the middle, not the creators of the work in question. Hell, the publishers are licking their chops as to how many artists they can fire and replace with AI, as if all their whinging about Napster and KaZaA 24 years ago was just a puff piece.

Re: Apple's On-Device and Server Foundation Models

#86
post #5

Earlier quoted context omitted.

So built on stolen data essentially.

Does that imply I just stole your comment by reading it? No snark intended; I’m seriously asking. If the answer is “no” then where do you draw the line?

I think scale is what changes the nature of the thing. At the point where you're having a machine consume billions of documents I don't think you could reasonably call that reading anymore. But what you are doing in my eyes is indexing, and the legal basis for that is heavily dependent on what you do with it.

If a human reads it that would be a reproduction of the work, but if you serve that page as a cache to a human you're okay, usually.

If you compile all that information in a database and use it to answer search queries that's also okay, and nothing forbids you from using machine learning on that data to better answer those search queries.

Both of the above are actually being challenged right now but for the time being they're fine.

But that database is a derivative work, in that it contains copyrighted material and so how you use it matters if you want to avoid infringement — for example a Google employee SSHing to a server to read NYT articles isn't kosher.

What isn't clear is whether the model is a derivative work. Does it contain the information or is it new information created from the training data Sure, if you're clever you could probably encode information in the weights and use it as a fancy zip file but that's a matter of intent. If you use Rewind or Windows Recall and it captures a screenshot of a NYT article and then displays it back to you later is that a reproduction? Surely not. And that's an autonomous system that stores copywritten data and regurgitates it verbatim.

So if it's impractical to actually use it for piracy and it very obviously isn't anyone's intent for it to be used as such then I think it's hard to argue it shouldn't be allowed, even on data that was acquired through back channels.

But copyright is more political than logical so who knows what the legal landscape will be in 5 years, especially when AI companies have every incentive to use their lawyers to pull the ladder up behind them.

Re: Apple's On-Device and Server Foundation Models

#87
post #81
post #71

Earlier quoted context omitted.

Where did you get the idea that's its way better than openai's? Aren't they both proprietary?

Apple isn’t collecting data from their customers. Edit: to feed back into their AI training.

[flagged]

Re: Apple's On-Device and Server Foundation Models

#88
post #55
post #44

Earlier quoted context omitted.

> There is no AppleBot-Extended. And if you blocked it in the past it remains blocked. You said there is no Applebot-Extended. The link says otherwise.

It's still true that there's no Applebot-Extended if it isn't crawling pages. Rather it's a marker to ask Applebot to limit what it does with your pages.

Isn't it still true that if people wanted to have their website show up in search in the past (so they didn't block Applebot), then it's too late to mark it as "no training" now, since it's already been scraped?

I guess it can be useful for data published in the future.

Re: Apple's On-Device and Server Foundation Models

#89

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

Piracy requires multiparty conspiracy against the establishment. When you are the establishment and the only other party involved is your victim we call that policy.

Re: Apple's On-Device and Server Foundation Models

#90
post #40

Earlier quoted context omitted.

Apple just did more to make this a privacy focused feature versus just a data mine than literally anyone else to date and still people complain. Public content on the internet is public content on the internet - I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.

Public content is still subject to copyright, and I doubt that AppleBot only scrapes content carrying a suitable license. And "fair use" (which is unclear if it applies), in case you want to invoke it, is a notion limited to the US and only a handful of other countries.

All you have to do is drop a token swear word into your content and they remove it from the dataset. Easy.
Post reply on HN