Earlier quoted context omitted.
It's a perfect reason to get consumers to buy a new phone every year again! They got bored of the camera.
Right now the models are doubling in performance (by the METR time horizon metric at least) every 4 months, so 3 doublings in a year; conversely, I hear (not my field) it takes around a year to make a prototype IC and another year to turn that into mass production, i.e. if the next (late-2026 model) iPhone has a chip like this, it will likely be with, at best, a late-2024 set of weights. I think you can get open-weig…
AMD acquires Taalas to boost inference performance by etching models in silicon
571–580 of 712 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#572I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…
If you make chips, you want models to be free.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#573Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#574Earlier quoted context omitted.
Right now the models are doubling in performance (by the METR time horizon metric at least) every 4 months, so 3 doublings in a year; conversely, I hear (not my field) it takes around a year to make a prototype IC and another year to turn that into mass production, i.e. if the next (late-2026 model) iPhone has a chip like this, it will likely be with, at best, a late-2024 set of weights. I think you can get open-weig…
wait, analog circuits? can you elaborate this?
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#575Earlier quoted context omitted.
Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
I'd gladly pay for a Claude Opus 4.6 Thinking High in silicon and use it for 1-2 years. It's good enough for many coding tasks.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#576TBH Taalas was a company I was existed about as a consumer. A dense model like gemma4-31b or qwen3.6-27b running at 10k t/s sounds like an awesome thing to have. Would be willing to pay GPU prices for it.
My 7900XT runs it at 35 tokens per second.
Taalas HC1 was clocked at 17000 tokens/s.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#577Economic and financial ripple effects would be huge aside from the obvious:
- reduction in electricity usage
- OpenAI / Anthropic are dead in the water unless they start to license their models to fabs.
- Every single one of those GPUs that all of those massive data centers contain become paperweights.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#578Big if it pans out. Economic and financial ripple effects would be huge aside from the obvious: - reduction in electricity usage - OpenAI / Anthropic are dead in the water unless they start to license their models to fabs. - Every single one of those GPUs that all of those massive data centers contain become paperweights.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#579Earlier quoted context omitted.
Right now the models are doubling in performance (by the METR time horizon metric at least) every 4 months, so 3 doublings in a year; conversely, I hear (not my field) it takes around a year to make a prototype IC and another year to turn that into mass production, i.e. if the next (late-2026 model) iPhone has a chip like this, it will likely be with, at best, a late-2024 set of weights. I think you can get open-weig…
wait, analog circuits? can you elaborate this?
The reason we don't do this in general (any more) is that for long chains between input and output it has been much too difficult to avoid accumulation of errors. LLMs happen to be extremely resilient to errors like this, which is also why we can use e.g. 4-bit weights.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#580Earlier quoted context omitted.
I’m on the IOS 27 beta and Siri did those two tasks (weather/phone) flawlessly. It’s a lot better than it used to be.
Thanks for trying that! Very interesting. Can you say this to it: "Hey Siri [wait for it to come up] - please send me an email with the temperature right now so I have it for my records." and see if it can complete the task without any backtalk or misunderstanding, and if you get exactly what you asked for. (It's a really clear request.) Should be 1 statement, no clarification, conversation, random search results, ("…
Subject: Current Temperature Body: The current temperature is 27°C in .