Live data from Hacker News

Apple's On-Device and Server Foundation Models

machinelearning.apple.com

71–80 of 562 posts

Re: Apple's On-Device and Server Foundation Models

#71

Earlier quoted context omitted.

Apple just did more to make this a privacy focused feature versus just a data mine than literally anyone else to date and still people complain. Public content on the internet is public content on the internet - I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.

Oh no, don't get me wrong. I like the privacy features, it's already way better than OpenAI's "we make it proprietary so we can spy on you" approach. What I don't like is the hypocrisy that basically every AI company has engaged in, where copying my shit is OK but copying theirs is not. The Internet is not public domain, as much as Eric Bauman and every AI research team would say otherwise. Even if you don't like cop…

Where did you get the idea that's its way better than openai's? Aren't they both proprietary?

Re: Apple's On-Device and Server Foundation Models

#72

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

Apple just did more to make this a privacy focused feature versus just a data mine than literally anyone else to date and still people complain. Public content on the internet is public content on the internet - I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.

Is that how copyright works now? I didn't see that they'd changed that law.

Re: Apple's On-Device and Server Foundation Models

#74
post #71

Earlier quoted context omitted.

Oh no, don't get me wrong. I like the privacy features, it's already way better than OpenAI's "we make it proprietary so we can spy on you" approach. What I don't like is the hypocrisy that basically every AI company has engaged in, where copying my shit is OK but copying theirs is not. The Internet is not public domain, as much as Eric Bauman and every AI research team would say otherwise. Even if you don't like cop…

Where did you get the idea that's its way better than openai's? Aren't they both proprietary?

Without the "so we can spy on you" part.

Re: Apple's On-Device and Server Foundation Models

#75
post #20

Earlier quoted context omitted.

I don’t actually think this is complicated and reading a comment is not the same thing as scraping the internet and you obviously know that. A few factors that come to mind would be: - scale - informed consent which there was none in this case - how you are going to use that data. For example using everybody others work so the worlds richest company can make more money from it while giving back nothing in return is a…

I think it's even simpler than that: incentives. The entire premise of copyright law (and all IP law) is to protect the incentive to create new stuff, which is often a very risky and highly time or capital intensive endeavor. So here's the question: Does a person reading a comment destroy the incentive for the author to post it? No. In fact, it is the only thing that produces the incentive for someone to post. People…

Thanks, this is a helpful comment.

It isn’t clear to me that these models destroy incentive to create. I mean, ChatGPT can generate comments in my style all day, and yet I’m still incentivized to comment.

I fancy myself a photographer. I still want to take photos even if DALL-E 4 will generate better ones.

What even is the point of creating art? I think there are two purposes: personal expression and enjoyment for others.

People will continue to express themselves even if a bot can produce better art.

And if a bot can produce enjoyment for others en masse, then that seems like a huge win for everybody.

Re: Apple's On-Device and Server Foundation Models

#76
post #49

Earlier quoted context omitted.

How are people supposed to block it when they stole all the data first and then only after that point they decide to even tell anyone what user agent they need to block and how they are planning to exploit your work for their profit.

You just have a rule that says block everything except crawlers: A, B, C. Also the AppleBot was known about before it appeared in Siri.

So you expect all websites to block FoobarSearch so it never gets off the ground and becomes a big search engine that people know to unblock.

Then FoobarSearch learns to ignore robots.txt wildcards, and we're back at square one.

IIRC this happened to DDG or Bing.

Re: Apple's On-Device and Server Foundation Models

#77
post #8

> Our foundation models are fine-tuned for users’ everyday activities, and can dynamically specialize themselves on-the-fly for the task at hand. We utilize adapters, small neural network modules that can be plugged into various layers of the pre-trained model, to fine-tune our models for specific tasks. For our models we adapt the attention matrices, the attention projection matrix, and the fully connected layers in…

I think it is just LoRA, you can call the LoRA weights as adapters

Re: Apple's On-Device and Server Foundation Models

#78

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

Apple just did more to make this a privacy focused feature versus just a data mine than literally anyone else to date and still people complain. Public content on the internet is public content on the internet - I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.

> I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.

Until LLMs came along, most large-scale internet scraping was for search engines. Websites benefited from this arrangement because search engines directed users to those websites.

LLMs abused this arrangement to scrape content into a local database, compress that into a language model, and then serve the content directly to the user without directing the user to the website.

It might've been legal, but that doesn't mean it was ethical.

Re: Apple's On-Device and Server Foundation Models

#79
post #20

Earlier quoted context omitted.

I don’t actually think this is complicated and reading a comment is not the same thing as scraping the internet and you obviously know that. A few factors that come to mind would be: - scale - informed consent which there was none in this case - how you are going to use that data. For example using everybody others work so the worlds richest company can make more money from it while giving back nothing in return is a…

I personally disagree but you make fair points. Scale: Many companies (e.g. Google, Bing) have been scraping at scale for decades without issue. Why does scale become an issue when an LLM is thrown into the mix? Informed consent: I’m not sure I fully understand this point, but I’d say most people posting content on the public internet are generally aware that people and bots might view it. I guess you think it’s diff…

There's a big difference between scraping a website so you can direct curious people to it (Googlebot) and scraping a website so you can set up a new website that conveys the same information, but earns you money and doesn't even credit the sources used (which these LLM services often do).

There is a whole genre of copyright infringement where someone will scrape a website and create a per-pixel copy of it but loaded up with ads, and blackhat SEOed to show up above the original website on searches. That's bad, and to the extent that LLMs are doing similar things, they are bad too.

Imagine I scrape your elaborate GameFAQs walkthrough of A Link to the Past. I could 1) use what I learn to direct curious people to its URL, or 2) remove your name from it, cut it into pieces, and rehost the content on my own page, mashed up with other walkthroughs of the same game. Then I sell this service as a revolutionary breakthrough that will free people from relying on carefully poring through GameFAQs walkthroughs ever again.

People will get mad about the second one, and to the extent what LLMs do is like that, will get mad at LLMs.

Re: Apple's On-Device and Server Foundation Models

#80
post #5

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

So built on stolen data essentially.

What’s the problem with that? Reproducing copyrighted works in full is problematic obviously. But if I learned English by watching American movies, I didn’t steal the language from the movie studios, I learned it.
Post reply on HN