Live data from Hacker News

Apple Foundation Models

platform.claude.com

171–180 of 244 posts

Re: Apple Foundation Models

#171
post #95

Earlier quoted context omitted.

I use both Claude and Codex and don’t see any meaningful difference between the two. My use case is modeling semi complex physical processes (energy and manufacturing) in code for simulations. I also have to do a good fair of automation via scripting in Python or PowerShell for manipulating data as well as legacy code analysis (C, Fortran, COBOL). Given I provide the models with the information and documentation they…

Did you find much of a difference between Fable and Opus?

Yes. Fable is much more organized and consistent at taking small bites of the (sorry) apple when solving a problem. Specifically I'm talking about a machine learning problem I'd been working on for awhile with Opus and it was (and is, again) constantly stating that all the signal is exploited, everything is now overfit, etc, etc, etc. The first day I pointed Fable at the situation I got a 10% improvement by paying attention to the little details that Opus instead took slightly negative results and extrapolated to "fully exploited". I've had to drop back, again, to forcing Opus to explain what it's looked at and the detail it has quietly assumed away.

It's like the difference to talking to two smartest kids in a class, but one really belongs a grade higher - and the other hasn't learned yet to ask the questions that encourage it to dig in that little bit more for the additional multi-order effects.

Re: Apple Foundation Models

#172

This is Apple commoditizing LLMs while keeping control of the UX. They are a hardware company and will keep selling the best machine for AI use. Well done.

Does “the best machine for AI use” apply here considering these models are still server-side?

Apple's been trying to make the marketing appeal that "Private Compute Cloud" is also a hardware project. Given it seems to rely on low level details of device Hardware Security Modules, it's maybe even at least a little bit more than just "marketing spin".

Re: Apple Foundation Models

#173

Earlier quoted context omitted.

I've found most of the frontier coding models require somewhere between 300GB to 1TB to run with full capabilities.

If only we could buy 1TB of unified memory in a Mac for $1k-$2k in total hardware costs. Apple would basically be able to extinguish the entirety of the market cap for Nvidia, OpenAI, Anthropic, and others all at once. In 10 years, I hope my MacBook Pro can run today's frontier models and has 1TB of unified Memory.

The Nvidia GB300 DGX Station, which isn't even going to hit 1TB total memory, is expected to launch at almost $100k. Bit of a pipe dream with memory prices where they're at.

Re: Apple Foundation Models

#174

Earlier quoted context omitted.

> Also missing from these discussions are e.g. Qwen, which is at least as good as one back from OpenAI or Anthropic’s frontiers. They're missing in the discussion because the ones you can run locally, aren't actually "one step away from other closed-source labs" in practice when you use them. They might benchmark as such, but they're sadly far away from measuring up to those scores except for very specific use cases,…

> the ones you can run locally, aren't actually "one step away from other closed-source labs" And they probably won’t be for at least another decade. Comparing like with like, flagship model running on the best hardware it can run on, Qwen is close.

> Qwen is close

I wish so badly this was true, but sadly today it just isn't.

Re: Apple Foundation Models

#176
post #97

Earlier quoted context omitted.

I think you're taking the written words a bit too literally here. Read it with a more lax filter and less literal word-meaning, and I think the original comment will become a bit clearer.

You know what, I've been a bit too snipe-y in my previous comments, and it led to to discussion devolving in unproductive ways. I'd genuinely like to understand where you're coming from more. I think we're all in agreement that this framework is very much about letting developers swap the models easily, and treat them as commodities. That seems pretty obvious. I do however still don't see how this has anything to do…

Thanks for the patience!

The way I see it, isn't about what is immediately there right now today, but what intent it signals, or what path Apple is planning. Yes, today it's ClaudeForFoundationModels, but the FoundationModels stuff will be used to allowed switching between models, probably without users noticing, and who knows what Apple will ultimately surface to users, tends to be in the direction of less user-control.

But there is a lot of assumptions, guesses and extrapolation from that, I think you're right if you focus only what's there right now, rather than trying to "see into the future" which harrouet basically started doing with their root comment.

Re: Apple Foundation Models

#177
post #171

Earlier quoted context omitted.

Did you find much of a difference between Fable and Opus?

Yes. Fable is much more organized and consistent at taking small bites of the (sorry) apple when solving a problem. Specifically I'm talking about a machine learning problem I'd been working on for awhile with Opus and it was (and is, again) constantly stating that all the signal is exploited, everything is now overfit, etc, etc, etc. The first day I pointed Fable at the situation I got a 10% improvement by paying at…

Had a very similar experience. Opus went "look, t-sne shows your features are neatly clustered" (it didn't) and left it at that. Fable didn't fully explore the problem/data, but it did go much further, implementing models to check for correlations and adjust feature clusters. Opus was able to finish the job after Fable was cut, but required much prodding (doing exactly what you described: pointing it towards things that look off and asking it, are you sure that's all there is to this?).

Re: Apple Foundation Models

#178
post #79
post #72

Earlier quoted context omitted.

Benedict Evans may be right after all; frontier models look more and more like telecom companies in the 90s. Billions and billions of investment in infrastructure while others further up the stack captured all the value.

In spite of their deeper pockets, massive datacenters, colosal amounts of user data, and hundreds of thousands of top developers, even Amazon, Meta, Microsoft, and Google are well behind. I think Evans is completely wrong. There are only 2 truly frontier models. (at least for now). And Anthropic seems to be leaving OpenAI behind so there might be only 1 in the near future. (which is scary/dangerous)

Is Google behind? The general opinions I read suggest Gemini is very competitive with Anthropic and OpenAI's top models.

Re: Apple Foundation Models

#180

Earlier quoted context omitted.

> the ones you can run locally, aren't actually "one step away from other closed-source labs" And they probably won’t be for at least another decade. Comparing like with like, flagship model running on the best hardware it can run on, Qwen is close.

> Qwen is close I wish so badly this was true, but sadly today it just isn't.

To be clear, I’m relaying my subjective experience comparing Opus and Qwen.
Post reply on HN