Earlier quoted context omitted.
Woa, good catch! Maybe they're doing better about at least being concrete about it, though I still have to side-eye "Users control their devices" (Even with root on macbooks I don't have access to everything running on it). However, the section that promises to open-source the cloud software are impressive and if true gives them more credibility than I assumed. I would still look out for places where devices they do…
> Even with root on macbooks I don't have access to everything running on it Just disable System Integrity Protection and then you do.
Apple's On-Device and Server Foundation Models
111–120 of 562 posts
Re: Apple's On-Device and Server Foundation Models
#112Earlier quoted context omitted.
I personally disagree but you make fair points. Scale: Many companies (e.g. Google, Bing) have been scraping at scale for decades without issue. Why does scale become an issue when an LLM is thrown into the mix? Informed consent: I’m not sure I fully understand this point, but I’d say most people posting content on the public internet are generally aware that people and bots might view it. I guess you think it’s diff…
There's a big difference between scraping a website so you can direct curious people to it (Googlebot) and scraping a website so you can set up a new website that conveys the same information, but earns you money and doesn't even credit the sources used (which these LLM services often do). There is a whole genre of copyright infringement where someone will scrape a website and create a per-pixel copy of it but loaded…
"Crediting the sources used" is not really a principle in copyright law. (Funny enough, online fanartists seem determined to convince everyone it is as a way of shaming people into doing it.)
Whether or not a use is transformative is protective though, and is what both of those cases rely on.
Re: Apple's On-Device and Server Foundation Models
#113> Our foundation models are trained on Apple's AXLearn framework, an open-source project we released in 2023. It builds on top of JAX and XLA, and allows us to train the models with high efficiency and scalability on various training hardware and cloud platforms, including TPUs and both cloud and on-premise GPUs. Interesting that they’re using TPUs for training, in addition to GPUs. Is it both a technical decision (J…
Re: Apple's On-Device and Server Foundation Models
#114Why isn't there a comparison with the Llama3 8b in the "benchmarks" ?
I believe it is because llama 3 8B beats it, which would make it look bad. The phi-3-mini version they used is the 4k which is 3.8B, while LLama 3 8B would be more comparable to phi-3 small (7B) which also considerably better than phi-3-mini. Likely both phi-3 small and llama 3 8B had too good results in comparison to Apple's to be added, since they did add other 7B models for comparison, but only when they won.
Re: Apple's On-Device and Server Foundation Models
#115Why isn't there a comparison with the Llama3 8b in the "benchmarks" ?
Re: Apple's On-Device and Server Foundation Models
#116Why isn't there a comparison with the Llama3 8b in the "benchmarks" ?
"If, on the Meta Llama 3 version release date, the monthly active users of the products or services made available by or for Licensee, or Licensee’s affiliates, is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta, which Meta may grant to you in its sole discretion, and you are not authorized to exercise any of the rights under this Agreement unless or until Meta otherwise expressly grants you such rights."
IANAL but my read of this is that Apple's not allowed to use Llama 3 at all, for any purposes, including comparisons.
Re: Apple's On-Device and Server Foundation Models
#117Earlier quoted context omitted.
Apple just did more to make this a privacy focused feature versus just a data mine than literally anyone else to date and still people complain. Public content on the internet is public content on the internet - I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet.
> I thought we had all agreed years ago that if you didn’t want your content copied, don’t make it freely available and unlicensed on the internet. Until LLMs came along, most large-scale internet scraping was for search engines. Websites benefited from this arrangement because search engines directed users to those websites. LLMs abused this arrangement to scrape content into a local database, compress that into a l…
Re: Apple's On-Device and Server Foundation Models
#118Earlier quoted context omitted.
Where did you get the idea that's its way better than openai's? Aren't they both proprietary?
Without the "so we can spy on you" part.
There's your bleeding, sorry truth there. It's only a matter of time until we get another headline like it.
Re: Apple's On-Device and Server Foundation Models
#119Earlier quoted context omitted.
Thanks, this is a helpful comment. It isn’t clear to me that these models destroy incentive to create. I mean, ChatGPT can generate comments in my style all day, and yet I’m still incentivized to comment. I fancy myself a photographer. I still want to take photos even if DALL-E 4 will generate better ones. What even is the point of creating art? I think there are two purposes: personal expression and enjoyment for ot…
It cheapens the incentive greatly. And you probably aren't selling your photos to make a living.
Re: Apple's On-Device and Server Foundation Models
#120Earlier quoted context omitted.
There's a big difference between scraping a website so you can direct curious people to it (Googlebot) and scraping a website so you can set up a new website that conveys the same information, but earns you money and doesn't even credit the sources used (which these LLM services often do). There is a whole genre of copyright infringement where someone will scrape a website and create a per-pixel copy of it but loaded…
> There's a big difference between scraping a website so you can direct curious people to it (Googlebot) and scraping a website so you can set up a new website that conveys the same information, but earns you money and doesn't even credit the sources used (which these LLM services often do). "Crediting the sources used" is not really a principle in copyright law. (Funny enough, online fanartists seem determined to co…
Legality aside, there is something very strange about a device that both 1) relies on your content to exist and could not work without it and 2) is attempting to replace it with its own proprietary chat interface. Googlebot mostly doesn't act like it's going to replace the internet, but Gemini and ChatGPT etc all are.
They're announcing "hi we are going to scrape all your data, put it into a pot, sell that pot back to you, and by the way, we are pushing this as a replacement for search, so from now on your only audience will be scrapers; all the human eyeballs will be on our website, which as we said before, relies on your work to exist."