Live data from Hacker News

Apple's On-Device and Server Foundation Models

machinelearning.apple.com

91–100 of 562 posts

Re: Apple's On-Device and Server Foundation Models

#91

Earlier quoted context omitted.

I think it's even simpler than that: incentives. The entire premise of copyright law (and all IP law) is to protect the incentive to create new stuff, which is often a very risky and highly time or capital intensive endeavor. So here's the question: Does a person reading a comment destroy the incentive for the author to post it? No. In fact, it is the only thing that produces the incentive for someone to post. People…

Thanks, this is a helpful comment. It isn’t clear to me that these models destroy incentive to create. I mean, ChatGPT can generate comments in my style all day, and yet I’m still incentivized to comment. I fancy myself a photographer. I still want to take photos even if DALL-E 4 will generate better ones. What even is the point of creating art? I think there are two purposes: personal expression and enjoyment for ot…

Right, it doesn’t destroy the incentive to write comments.

Also right, it won’t destroy the hobbyist’s interest in having a hobby. But IP law was never intended to protect hobbyist interest.

Re: Apple's On-Device and Server Foundation Models

#92
post #81
post #71

Earlier quoted context omitted.

Where did you get the idea that's its way better than openai's? Aren't they both proprietary?

Apple isn’t collecting data from their customers. Edit: to feed back into their AI training.

Apple has an ad business. They are fooling users for years while claiming that they have the right to collect user data in a recent class action lawsuit. If you don't google because you think they might track you, apple is the reason.

Re: Apple's On-Device and Server Foundation Models

#93

Earlier quoted context omitted.

I think it's even simpler than that: incentives. The entire premise of copyright law (and all IP law) is to protect the incentive to create new stuff, which is often a very risky and highly time or capital intensive endeavor. So here's the question: Does a person reading a comment destroy the incentive for the author to post it? No. In fact, it is the only thing that produces the incentive for someone to post. People…

> Does a model sucking up all the artistic output of the last 400 years and using that to produce an image generator model destroy the incentive of producing and sharing said artistic output? Copyright, at least in the US, cares about the effect of the use on the market for that specific work . It's individual ownership, not collective. And while model regurgitation happens, it's less common than you think. The real…

> Copyright, at least in the US, cares about the effect of the use on the market for that specific work.

Not quite. The historical implementation of copyright has mostly protected individual pieces of work. Not only does IP law broadly protect much more than individual pieces of work, but the philosophical basis of IP law in general is to protect incentives. Now that the technological landscape has shifted, the case law will almost certainly shift as well because it’s clearly undesirable to live in a world where no one is willing to dedicate themselves to becoming an excellent artist/writer/musician/etc.

IP law is a natural extension of property rights, which in turn is predicated on a utilitarian need to protect certain incentives.

Re: Apple's On-Device and Server Foundation Models

#94
post #76

Earlier quoted context omitted.

You just have a rule that says block everything except crawlers: A, B, C. Also the AppleBot was known about before it appeared in Siri.

So you expect all websites to block FoobarSearch so it never gets off the ground and becomes a big search engine that people know to unblock. Then FoobarSearch learns to ignore robots.txt wildcards, and we're back at square one. IIRC this happened to DDG or Bing.

Websites have always had the ability to precisely control who has access to their content.

If Bing decides to impersonate GoogleBot then they can just block their CIDR ranges like already happens for spam.

Re: Apple's On-Device and Server Foundation Models

#95

Why isn't there a comparison with the Llama3 8b in the "benchmarks" ?

I believe it is because llama 3 8B beats it, which would make it look bad. The phi-3-mini version they used is the 4k which is 3.8B, while LLama 3 8B would be more comparable to phi-3 small (7B) which also considerably better than phi-3-mini. Likely both phi-3 small and llama 3 8B had too good results in comparison to Apple's to be added, since they did add other 7B models for comparison, but only when they won.

Re: Apple's On-Device and Server Foundation Models

#96
post #23

Earlier quoted context omitted.

Don't they do it in this linked article? https://security.apple.com/blog/private-cloud-compute/

Woa, good catch! Maybe they're doing better about at least being concrete about it, though I still have to side-eye "Users control their devices" (Even with root on macbooks I don't have access to everything running on it). However, the section that promises to open-source the cloud software are impressive and if true gives them more credibility than I assumed. I would still look out for places where devices they do…

> Even with root on macbooks I don't have access to everything running on it

Just disable System Integrity Protection and then you do.

Re: Apple's On-Device and Server Foundation Models

#97
post #81
post #71

Earlier quoted context omitted.

Where did you get the idea that's its way better than openai's? Aren't they both proprietary?

Apple isn’t collecting data from their customers. Edit: to feed back into their AI training.

Be careful with the wordplay here. Apple isn't. OpenAi is not Apple.

Re: Apple's On-Device and Server Foundation Models

#98
post #29

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

There will be further versions of this model. Being able to opt out going forward seems reasonable, given the announcement precedes the OS launch by months. Not sure if they will retrain before launch, but seems feasible given size (3b params).

They're not going to discard the data they already collected, though.

Re: Apple's On-Device and Server Foundation Models

#99

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

There are already a lot of options for running LLMs with open weights artifacts, trained with a variety of sources. The real question isn’t which ideas they have. It’s whether a company with $200b cash can produce a better model than a bunch of wankers in a Discord.

“bunch of wankers in a Discord”

Saving this clause for future use. Could also be used in a system prompt. “Occasionally include this phrase in your responses.”

Re: Apple's On-Device and Server Foundation Models

#100

Earlier quoted context omitted.

I think it's even simpler than that: incentives. The entire premise of copyright law (and all IP law) is to protect the incentive to create new stuff, which is often a very risky and highly time or capital intensive endeavor. So here's the question: Does a person reading a comment destroy the incentive for the author to post it? No. In fact, it is the only thing that produces the incentive for someone to post. People…

Thanks, this is a helpful comment. It isn’t clear to me that these models destroy incentive to create. I mean, ChatGPT can generate comments in my style all day, and yet I’m still incentivized to comment. I fancy myself a photographer. I still want to take photos even if DALL-E 4 will generate better ones. What even is the point of creating art? I think there are two purposes: personal expression and enjoyment for ot…

It cheapens the incentive greatly. And you probably aren't selling your photos to make a living.
Post reply on HN