I'm surprised there's not more discussion about potential inflection points here. When technology gets faster, it opens up whole new classes of UX that were hard to predict For example, faster internet didn't mean being able to view 100x as many HTML4 web pages. It brought SaaS, streaming media and interactivity. I'm not good at predicting, but some ideas: 1. All information gets augmented in real time with personali…
AMD acquires Taalas to boost inference performance by etching models in silicon
501–510 of 712 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#502Earlier quoted context omitted.
It barely works even today, like Siri is laughably bad.
I mean the examples he gave definitely work. Mostly well I'd say as they are pretty primitive. What Siri is missing is more logical solutions and answers for recipes, etc (still suck even with chatgpt integration).
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#503Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#504Earlier quoted context omitted.
Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.
Not sure. You can fix the transistors but leave the connections between them open for flexibility, so you only need to change the manufacturing process for the upper masks for every new model.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#505Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#506I'm surprised there's not more discussion about potential inflection points here. When technology gets faster, it opens up whole new classes of UX that were hard to predict For example, faster internet didn't mean being able to view 100x as many HTML4 web pages. It brought SaaS, streaming media and interactivity. I'm not good at predicting, but some ideas: 1. All information gets augmented in real time with personali…
Edit: also consider centralized Room-641A-type surveillance when models summarize and/or flag all calls processed by public telephony
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#507I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
They are: https://openai.com/index/cerebras-partnership/ My guess is they only consider Luna "good enough" to justify the immense up-front investment to put it onto silicon, but Luna at 10x the current speed would be killer. If they're really pursuing live voice conversations with a hardware assistant, latency is more important than accuracy (for complex questions the assistant could always say something like "wait a…
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#508This is neat but IMO a little crazy. Something I personally haven’t seen much of, in all the discussions of model benchmarks and AI breakthroughs, is a distinction between “peak performance” and “reliable performance”. The “peak performance” of frontier models is very high: they’re solving open math problems, analyzing large codebases, etc. But my subjective impression is that “reliable performance” is mid at best: o…
Now, that’s slow and expensive although seems to work quite well (haven’t really evaled this properly, don’t have the time). If inference can be made fast and cheap, multi-model approaches like this would become more viable for more applications.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#509This move undercuts NVIDIA directly.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#510Earlier quoted context omitted.
I think the lifecycle for these chips could stretch far longer. If you're offering these models on a two year lifecycle, then you'd be able to stand up your top tier (wouldn't need to be frontier) at high speed. Run (for example) Kimi K3 on it and give it a brand name: AcmeAI Carbon Market it as your premier (only) model at high throughput. Two years later you stand up MSICs for the new state of the art with entirely…
I'm somewhat doubtful that we will be seeing something as large as Kimi K3 in silicon any time soon. This tech can definitely scale up from the current 8B prototype, but - at least as far as my limited understanding of the tech involved goes - you cannot just ASIC a trillion weights model due to physical size constraints. ___ Specification HC1 Model Llama 3.1 8B (hardwired) Process TSMC 6nm Die size 815mm² ___ So the…