Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

651–660 of 710 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#651
post #586

Earlier quoted context omitted.

Scaling a model on a chip is quite hard. ChatJimmy is based on Llama 3.1 8 billion. Kimi K3 has 2.8 trillion parameters. That are 350x more parameters. I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or maybe even Gemma 4 26 billion A4B at some point is running on a chip.

350x is only about 10-20 years of improvement, using CPU FLOPS as the benchmark.

wouldnt you want gpu, fpga, or dsp as the benchmark?

its lots of parallel calculations, rather than one blazing fast one

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#652
post #611
post #388

Earlier quoted context omitted.

Errors compound, and making 1000 wrong decisions per hour, will not result in something useful. Maybe you‘ve tried setting up guardrails for good design or architecture at some point? I think it’s simply not possible to do that. It would certainly be an accelerator for people who know exactly what they want. And it would remove multi tasking, which I‘d appreciate.

I don't have a great answer but you pose a great question. Obviously a CTO is not going to walk away from the technology just because it's not good enough. That much more incentive for someone to create a powerful enough harness that can direct that power safely and productively. Like a nuclear core, we'll need to come up with the graphite rods and water tank. And if tokens are essentially free, why not , for every m…

I do spend 5x more tokens on planning and reviewing, than for implementation. But architecture is still nothing I can delegate.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#653

I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…

It's possible in the future we will have Rick and Morty style AI in literally everything just because it's so easy to add it.

    Sentient Switchblade: "Hi Beth! You've gotten taller! Shall we resume stabbing?"

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#655
post #636

I see a lot of discourse about it being fast-to-deprecation. But I see it a different way personally. Modern LLMs are trying to do more with less. Focus on doing the right thing the first time. Even if we squeeze dumb LLMs, the significantly faster speed means quicker iterations. So a bad decision doesn't cost the time and inference costs that it cost before. It theoretically changes the scale of errant token spend.…

Sounds to me like you would hire 20 barely-paid interns instead of 2 competent programmers.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#656
post #655
post #636

I see a lot of discourse about it being fast-to-deprecation. But I see it a different way personally. Modern LLMs are trying to do more with less. Focus on doing the right thing the first time. Even if we squeeze dumb LLMs, the significantly faster speed means quicker iterations. So a bad decision doesn't cost the time and inference costs that it cost before. It theoretically changes the scale of errant token spend.…

Sounds to me like you would hire 20 barely-paid interns instead of 2 competent programmers.

if it fits, it ships

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#657

I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…

Now the fun part, how will having an LLM in my washing machine help anything

Tell the washing machine. "Washlexa I'm putting my gym clothes in. It was a hard workout I sweated a lot!"

and the washer now knows what to do. No more temp, length, 2nd rinse proxy controls.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#658

Earlier quoted context omitted.

Now the fun part, how will having an LLM in my washing machine help anything

Tell the washing machine. "Washlexa I'm putting my gym clothes in. It was a hard workout I sweated a lot!" and the washer now knows what to do. No more temp, length, 2nd rinse proxy controls.

And result is tiny little clothes, because 80C seemed right for the job.

Washlexa:”Sorry, you were right I was supposed to use regular programme, but I used wrong one, do you want me to wash them again?”

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#659

Earlier quoted context omitted.

Personally I think Apple should have acquired them. if you could burn a gemma4 class model into an iphone and actually get extremely low latency and low battery usage it would feel like the future IMO. even if it means you wont get frontier intelligence, there might actually be incentive to buy a new mobile device every year again.

The Taalas chips are not physically small. And part of their secret (if you look at the design) is just locating a bunch of memory soldered on the edges ( I belive higher amounts of SRAM ? )

Sure but the first step to having something that is physically small, small enough to cram into an iPhone, is to have something that, at first, isn't.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#660

I think what will happen is what happened to something like 4K video decoding before where it ends up in silicon costing almost nothing to run extremely fast on device. "Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption. Probably this will also happen for software engineering. Some usb-power…

It's possible in the future we will have Rick and Morty style AI in literally everything just because it's so easy to add it. Sentient Switchblade: "Hi Beth! You've gotten taller! Shall we resume stabbing?"

Nightblood? is that you?
Post reply on HN