Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

591–600 of 712 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#591

Earlier quoted context omitted.

The investment in bigger machines at the fab might set you back billions. I don't know about the lithography technology either, how easy you can scale it to larger wafers?

You also need to worry about yields, Apple, AMD etc can sell ”bad” chips as lower core versions, if you’re depending on whole wafer you have little room for error.

Love the idea of discounts based on model error. ”this one doesn’t know what butterflies are, it’s on sale for 8% off”

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#592

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

I wish we lived in a reality where Framework was anywhere near rich enough to acquire them. I'd love to have models on a chip that I could swap in at a whim. That would do the opposite by eliminating moats.

I guess there's a tiny chance AMD makes something like that happen. It seems like a great way to get people and orgs to pay a few hundred bucks every 6 months or so.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#593

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

>Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

Google is no longer a serious player in frontier AI. I doubt they will ever hit a SOTA model again.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#594
post #554

Earlier quoted context omitted.

At some point it’s got to be good enough for the normal “phone stuff” that appeal to most users. So they wouldn’t suffer from FOMO because they didn’t wait for the next model. Every phone gimmick went through the same evolution curve until it passed the “good enough” point and eventually plateaued.

Yes, but irrelevant. While these models are improving at the present rate, the manufacturer can save money at no loss of feature-bullet-point-on-website by letting you download a model after you bought the thing and running it on normal hardware. The rate of change to the models has to be slower than the hardware roll-out to be worth a hardware solution. If "good enough" happens before then, that just means the user…

You might be right but hard to tell without analyzing costs and benefits. Is a cutting edge model for phone stuff worth the slower performance and battery drain for example?

The rate of change by itself doesn’t tell you the whole story because of costs and diminishing returns. So what if your model is twice as good if it’s 10x the cost and it saves you 1ms? Everything else about phones reached “good enough for a phone” levels in years, and then got minimal generational improvements.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#595
post #580

Earlier quoted context omitted.

Thanks for trying that! Very interesting. Can you say this to it: "Hey Siri [wait for it to come up] - please send me an email with the temperature right now so I have it for my records." and see if it can complete the task without any backtalk or misunderstanding, and if you get exactly what you asked for. (It's a really clear request.) Should be 1 statement, no clarification, conversation, random search results, ("…

It brings up a preview of the email and you have to tap or tell it to send it from there but otherwise it worked for this as well. Subject: Current Temperature Body: The current temperature is 27°C in .

thanks! useful.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#596
post #419

Earlier quoted context omitted.

I don’t think average user _needs_ to solve frontier challenges. ”Call to Jane”, ”turn on the lights” and ”what’s the weather this afternoon” is more like it I would guess. Ofc if the model has some critical bugs that’s another matter.

Your examples worked on phones for over a decade. Maybe baking in a model that is "certified" to have some unconditioned truths + rest is pulled from external models/store could make sense. But AFAIK that doesn't exist and I'm not sure it can possibly be made. Perhaps society as a whole at least can work on an open corpus of training data, but I'm not holding my breath on this.

I just asked Siri

"Hey Siri, what's the weather in tomorrow" and it gave me a town with the same name ~800 miles from me.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#597
post #256

Earlier quoted context omitted.

You could take your silicon chip and have it re-etched only with model diffs for an upgraded version.

How does that work, as in re-etching of silicon? Any pointers to read?

Someone will figure it out.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#598
post #121

Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.

We have some processes running on models released a year ago (which we're updating, but still)

The speed is incredible. It doesnt matter if you are ~30-300 days behind

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#599
post #400

Earlier quoted context omitted.

Their PoC chips are big, but then it's ridiculously fast (have you seen chatjimmy.ai?). Also they must be holding a bunch of patents.

Its a cool demo, but its gpt-3.5 level stupid, or worse. edit: Ok, I will self-apologize. Its apparently a 3B model. Mighty impressive for what it does.

How time flies. Just 3.5 years ago, gpt-3.5 was touted as almost AGI, we're all going to be replaced by machines and worst case they will kill us all. And here we are, not much later, and it serves as the benchmark for "stupid"..

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#600

Earlier quoted context omitted.

Scaling a model on a chip is quite hard. ChatJimmy is based on Llama 3.1 8 billion. Kimi K3 has 2.8 trillion parameters. That are 350x more parameters. I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or maybe even Gemma 4 26 billion A4B at some point is running on a chip.

It's quite telling that the 8B Taalas chip was already reticle-sized on TSMC N6. I mean, we're talking about a process that does ~100 MTr/mm², ROM needs about one transistor per bit, but can probably be packed more densely than general logic. Something like, say, 150 megabit/mm² is not a lot. N6 has a 850 mm² reticle limit. This roughly tracks, the article says the chip has 8B parameters and apparently spends about h…

> (each HBM3 die is >1000mm² of silicon)

Did you mean that each HBM3 stack is that large? Because it only takes one glance to see that the memory chips are much smaller than reticle-sized GPUs they sit next to.

Post reply on HN