Earlier quoted context omitted.
Your examples worked on phones for over a decade. Maybe baking in a model that is "certified" to have some unconditioned truths + rest is pulled from external models/store could make sense. But AFAIK that doesn't exist and I'm not sure it can possibly be made. Perhaps society as a whole at least can work on an open corpus of training data, but I'm not holding my breath on this.
>Your examples worked on phones for over a decade. Nope. And not only not a decade ago, right now. If you have an Android or iPhone, you can give it clear and easy to understand instructions that Gemma 4 could complete[1] if it had tool calls on it, and that 100.00% of Claude, ChatGPT, Grok, Kimi, you name it, could understand and all complete if they had the access. The phones will fail to complete it. I just tried…
AMD acquires Taalas to boost inference performance by etching models in silicon
551–560 of 712 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#552This is neat but IMO a little crazy. Something I personally haven’t seen much of, in all the discussions of model benchmarks and AI breakthroughs, is a distinction between “peak performance” and “reliable performance”. The “peak performance” of frontier models is very high: they’re solving open math problems, analyzing large codebases, etc. But my subjective impression is that “reliable performance” is mid at best: o…
What I imagine an on-device model should be doing is just translate natural language to search requests and calls to tools manipulating retrieved data - much like no model currently does calculations and instead they open up calculator and use that instead.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#553Earlier quoted context omitted.
It's a perfect reason to get consumers to buy a new phone every year again! They got bored of the camera.
Right now the models are doubling in performance (by the METR time horizon metric at least) every 4 months, so 3 doublings in a year; conversely, I hear (not my field) it takes around a year to make a prototype IC and another year to turn that into mass production, i.e. if the next (late-2026 model) iPhone has a chip like this, it will likely be with, at best, a late-2024 set of weights. I think you can get open-weig…
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#554Earlier quoted context omitted.
Right now the models are doubling in performance (by the METR time horizon metric at least) every 4 months, so 3 doublings in a year; conversely, I hear (not my field) it takes around a year to make a prototype IC and another year to turn that into mass production, i.e. if the next (late-2026 model) iPhone has a chip like this, it will likely be with, at best, a late-2024 set of weights. I think you can get open-weig…
At some point it’s got to be good enough for the normal “phone stuff” that appeal to most users. So they wouldn’t suffer from FOMO because they didn’t wait for the next model. Every phone gimmick went through the same evolution curve until it passed the “good enough” point and eventually plateaued.
The rate of change to the models has to be slower than the hardware roll-out to be worth a hardware solution. If "good enough" happens before then, that just means the user gets a software solution.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#555Earlier quoted context omitted.
>Your examples worked on phones for over a decade. Nope. And not only not a decade ago, right now. If you have an Android or iPhone, you can give it clear and easy to understand instructions that Gemma 4 could complete[1] if it had tool calls on it, and that 100.00% of Claude, ChatGPT, Grok, Kimi, you name it, could understand and all complete if they had the access. The phones will fail to complete it. I just tried…
I’m on the IOS 27 beta and Siri did those two tasks (weather/phone) flawlessly. It’s a lot better than it used to be.
Can you say this to it: "Hey Siri [wait for it to come up] - please send me an email with the temperature right now so I have it for my records." and see if it can complete the task without any backtalk or misunderstanding, and if you get exactly what you asked for. (It's a really clear request.) Should be 1 statement, no clarification, conversation, random search results, ("Here's what I found!"), etc.
A normal frontier model can do that - or Siri can do it if it is properly connected to Claude, ChatGPT, Gemini, Grok, or any other frontier AI - but previously it was never properly connected.
If it can do this task, I might have to look into this again. It counts as a success if it sends yourself any email with the current temperature and you actually get it (it can include whatever other text in the email), and a failure if it talks back, says "here's what I found", says it can't, asks you any question, sends you an email that doesn't actually contain the current temperature, just reads you the temperature and then asks if you want it to send an email, etc. Should be 1 shot.
let me know if it works!
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#556I'm surprised there's not more discussion about potential inflection points here. When technology gets faster, it opens up whole new classes of UX that were hard to predict For example, faster internet didn't mean being able to view 100x as many HTML4 web pages. It brought SaaS, streaming media and interactivity. I'm not good at predicting, but some ideas: 1. All information gets augmented in real time with personali…
https://youtu.be/7NfyZhV1dKM?t=53
Imagine this demo but the apps render in real time generating real code.
Obviously not valuable because we have OS today, but could be your companies "WorkOS"
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#557I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.
... you would still have a mediocre phone with half-assed barely working features driven by locked down proprietary software
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#558That is a downside to be sure, but from a pure business perspective, "that's not a bug -- it's a feature!"... from a pure business perspective it's the ability to sell and resell, to purchasing and re-purchasing customers, way into the future -- that is, recurring revenue from the perspective of the company being able to make those future recurring sales...
In the above case, that company is AMD...
(Also, on a related note, it would be interesting to see what open source / open hardware work has currently been done to offload LLM weights (and/or anything else that could be offloaded to silicon ASIC's) to FPGA's...)
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#559Earlier quoted context omitted.
The size of model we're talking about running doesn't need much if any dram.
The chatjimmy demo is using a model that needs 6-18GB of VRAM. That's not exactly trivial. I could see it being feasible to get a Qwen-3.6-27b type of model done on something like this. Qwen-3.6-27b at 18tok/s would be a game changer.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#560Earlier quoted context omitted.
Cerebras is literally the entire wafer, so it can't get bigger. So where is the jump from 30x to 100x coming from? Node improvements only yield like 10-20% gains these days...
Could we not just make bigger wafers, if the technology called for it?