Live data from Hacker News

Apple's On-Device and Server Foundation Models

machinelearning.apple.com

61–70 of 562 posts

Re: Apple's On-Device and Server Foundation Models

#63
post #20

Earlier quoted context omitted.

Does that imply I just stole your comment by reading it? No snark intended; I’m seriously asking. If the answer is “no” then where do you draw the line?

I don’t actually think this is complicated and reading a comment is not the same thing as scraping the internet and you obviously know that. A few factors that come to mind would be: - scale - informed consent which there was none in this case - how you are going to use that data. For example using everybody others work so the worlds richest company can make more money from it while giving back nothing in return is a…

I personally disagree but you make fair points.

Scale: Many companies (e.g. Google, Bing) have been scraping at scale for decades without issue. Why does scale become an issue when an LLM is thrown into the mix?

Informed consent: I’m not sure I fully understand this point, but I’d say most people posting content on the public internet are generally aware that people and bots might view it. I guess you think it’s different when the data is used for an LLM? But why?

Data usage: Same question as above.

I just don’t see how ingestion into an LLM is fundamentally different than the existing scraping processes that the internet is built on.

Re: Apple's On-Device and Server Foundation Models

#64

> We train our foundation models on licensed data, including data selected to enhance specific features, as well as publicly available data collected by our web-crawler, AppleBot. Web publishers have the option to opt out of the use of their web content for Apple Intelligence training with a data usage control. And, of course, nobody has known to opt-out by blocking AppleBot-Extended until after the announcement wher…

There are already a lot of options for running LLMs with open weights artifacts, trained with a variety of sources. The real question isn’t which ideas they have. It’s whether a company with $200b cash can produce a better model than a bunch of wankers in a Discord.

Re: Apple's On-Device and Server Foundation Models

#65

Earlier quoted context omitted.

Does that imply I just stole your comment by reading it? No snark intended; I’m seriously asking. If the answer is “no” then where do you draw the line?

Data gets either stolen or freed depending on whether the guy who copied it is someone you dislike or like. Personally, I think that Apple is giving the data more exposure which, as I've been informed many times here, is much more valuable than paying for the data.

The irony of "do it for the exposure" is that everyone who actually wants to pay you in exposure isn't actually going to do that, either because they aren't popular enough to measurably expose you, or because they're so popular that they don't want to share the limelight.

AI is a unique third case in which we have billions of creators and no idea who contributed what parts of the model or any specific outputs. So we can't pay in exposure, aside from a brutally long list of unwilling data subjects that will never be read by anyone. Some of the training data is being regurgitated unmodified and needs to be attributed in full, some of it is just informing a general understanding of grammar and is probably being used under fair use, and yet more might not even wind up having any appreciable effect on the model weights.

None of this matters because nobody actually agreed to be paid in exposure, nor was it ever in any AI company's intent - including Apple - to pay in exposure. Data is free purely because it would be extraordinarily inconvenient if anyone in this space had to pay.

And, for the record, this applies far wider than just image or text generators. Apple is almost surely not the worst offender in the space. For example: all that facial recognition tech your local law enforcement uses? That was trained on your Facebook photos.

Re: Apple's On-Device and Server Foundation Models

#66

Why isn't there a comparison with the Llama3 8b in the "benchmarks" ?

Maybe it’s too new for them to have had time to include it in their studies?

Phi-3-Mini, which is in the benchmarks, was released after Llama3 8b

Re: Apple's On-Device and Server Foundation Models

#68

Earlier quoted context omitted.

Maybe it’s too new for them to have had time to include it in their studies?

Phi-3-Mini, which is in the benchmarks, was released after Llama3 8b

Llama 3 8B is really really good. Maybe it makes Apples models look bad? Or it could be a licensing thing where Apple can’t use Llama 3 at all, even just for benchmarking and comparison.

The license for the Llama models was basically designed to stop Apple, Microsoft and Google from using it.

Re: Apple's On-Device and Server Foundation Models

#69
post #46
post #7

Earlier quoted context omitted.

> just trained a new OS development AI on every OS Apple has ever written. …is there publicly visible source code for every OS Apple has ever written?

Partially: https://github.com/apple-oss-distributions/distribution-macO... https://github.com/apple-oss-distributions/distribution-iOS I'm not sure how it all fits together but people have even made an open source distribution of the base of darwin, the underlying OS: https://github.com/PureDarwin/PureDarwin

Apple's FOSS releases are purely command-line userland tools and their kernel, all the frameworks and servers[0] that make the UI work like you'd expect are 100% proprietary.

FWIW Apple has also been on a decades-long track of purging GPL-licensed code from macOS and replacing it with either permissively-licensed or proprietary equivalents. So they're obligated to release even less than they used to.

[0] AppKit/WindowServer for MacOS and UIKit/Springboard/Backboard for everything else

Re: Apple's On-Device and Server Foundation Models

#70

Earlier quoted context omitted.

Apple just did more to make this a privacy focused feature versus just a data mine than literally anyone else to date and still people complain. Public content on the internet is public content on the internet - I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.

Oh no, don't get me wrong. I like the privacy features, it's already way better than OpenAI's "we make it proprietary so we can spy on you" approach. What I don't like is the hypocrisy that basically every AI company has engaged in, where copying my shit is OK but copying theirs is not. The Internet is not public domain, as much as Eric Bauman and every AI research team would say otherwise. Even if you don't like cop…

> then copyleft is useless as a tactic to get the proprietary world to bend the knee

I have bad news

Post reply on HN