Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

701–710 of 710 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#701
post #692

Earlier quoted context omitted.

What I'm saying is that Apple will use these type of models etched into chips, and they will do it because it drives obsolescence, so they can shorten the upgrade cycle. They will do it because they figure out it's good for them.

You've confused engineering compromise for malice, and reversed the purpose. For the model capabilities and inference power draw, what alternative do you see to a (at least mostly) fixed hardware model?

No, you still don't understand what I'm saying. Yes, ASICs make inference faster, but also makes the hardware obsolete faster, if it's embedded in a phone. That sounds like a negative, but Apple is going to turn it into a positive for their business and use it to speed up upgrade cycles as cameras are no longer a driving factor and cycles have been getting longer.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#702

Earlier quoted context omitted.

You also need to worry about yields, Apple, AMD etc can sell ”bad” chips as lower core versions, if you’re depending on whole wafer you have little room for error.

Love the idea of discounts based on model error. ”this one doesn’t know what butterflies are, it’s on sale for 8% off”

Yes, "this one is obsessed with the golden gate bridge" will be for real this time.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#703
post #406

Earlier quoted context omitted.

The investment in bigger machines at the fab might set you back billions. I don't know about the lithography technology either, how easy you can scale it to larger wafers?

we seem to be in a phase of spending trillions on the computer, so while it isn’t likely, it isn’t impossible

[deleted]

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#704

Earlier quoted context omitted.

Now the fun part, how will having an LLM in my washing machine help anything

Tell the washing machine. "Washlexa I'm putting my gym clothes in. It was a hard workout I sweated a lot!" and the washer now knows what to do. No more temp, length, 2nd rinse proxy controls.

I now picture people talking to their washing machine like the intro of american psycho.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#705
post #546

Earlier quoted context omitted.

An "Expert" is really just an unfortunate name for what amounts to a dense part of a sparse matrix and that's also an oversimplification. It doesn't actually specialise in anything in particular that one can point to. For this reason you can really transfer them between models.

Is it still true? I'd assume you should be able to freeze the matrix and unfreeze an expert block, before feeding particularly chosen training data. Or that doesn't work?

That's more or less the idea behind Low-Rank Adaptation, or LoRA.

There's also Mixture of LoRA Experts, which instead of slicing up the model and routing through that, routes through different LoRAs.

But it all comes with tradeoffs, as you have to train and run the gating network doing the routing, which also comes at a cost.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#706

Earlier quoted context omitted.

They are loaded on top in the sense that they are contained in the upper layer(s) of the chip. So when you want to change the weights, you have to produce fewer masks for fabrication, which reduces cost and time to market.

Oh... That's not great - it'd be nice if it had a way to push updates without building a new chip. OTOH, maybe because of this our future cyberdecks will have cartridge ports.

I believe the fully baked-in nature is pretty important for the perf they get. With a cartridge port you now have a bus between the weights and the compute and the weights and that bus can become a bottleneck.

So I think we're looking at a spectrum here:

- Fully fixed function - i.e. Taalas

- Fixed function transformer unit (or whatever other AI architecture) with a "cartridge" for weights etc - i.e. Etched

- Flexible TPU/GPU type stuff

I think the middle of the spectrum is a bit of a dead space ATM because new models have recently been coming with significant updates to the architecture (like MoE, MTP) so by the time you have new weights you wanna load, you also want to replace the compute too. So really you probably either say "I can tolerate an old model, but I want it fast as FUCK" and go for Taalas-style, or you say "I want a near-frontier model" and you have to use flexible compute anyway.

But, caveat: this comment seems to be making me sound more knowledgeable than I actually am. Take this with a grain of salt.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#707
post #385

Earlier quoted context omitted.

It's already kind of that way with MCP servers popping up everywhere. The JIRA MCP server is like a couple orders of magnitude faster to work with than the website itself.

That’s their API with extra steps, or am I missing something? That was always faster.

The extra steps with AI are a better and faster user experience than their website.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#708
post #191

Earlier quoted context omitted.

Ok, how long until nvidia gives us a new GPU?

I don't follow. How is that related? GPUs don't have fixed memory. You don't throw them away when you want to load a new model. NVIDIA will probably give us a new GPU when someone competent in the free market decides they want wheelbarrows full of money. Unfortunately, AMD is entirely, incomprehensibly, incompetent, to the point where I can only assume they're colluding with Nvidia, behind the scenes.

Point being that hardware generations can be very quick and as updates get made GPUs go out of date and can’t run the latest models. All types of hardware consumer of otherwise are always improving, which means that tying a model the hardware is not going to lead to increased obsolescence, any more than the hardware itself does. If an AI is general purpose then there is no problem with only having that one model baked into the chip, when the benefit is a huge speed and efficiency improvement over running it on a GPU.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#709

Earlier quoted context omitted.

Now the fun part, how will having an LLM in my washing machine help anything

Tell the washing machine. "Washlexa I'm putting my gym clothes in. It was a hard workout I sweated a lot!" and the washer now knows what to do. No more temp, length, 2nd rinse proxy controls.

make the "washing machine" a bucket with a robotic arm that you can put in any bathtub

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#710
post #400

Earlier quoted context omitted.

Its a cool demo, but its gpt-3.5 level stupid, or worse. edit: Ok, I will self-apologize. Its apparently a 3B model. Mighty impressive for what it does.

How time flies. Just 3.5 years ago, gpt-3.5 was touted as almost AGI, we're all going to be replaced by machines and worst case they will kill us all. And here we are, not much later, and it serves as the benchmark for "stupid"..

Yup.

Where I work we invited a bunch of people circa May 2023: a few top tier academics, a few government and NGO officials in charge of our industry, and a few startup founders. We are a big company, so people were kind enough to come and give speeches - and discuss.

This was a roundtable on what's going to happen. You could play some of the speeches verbatim today and they would not feel out of place. And that is telling. We had a Stanford professor saying that gpt-3.5 can do everything he can - just better, and he feels his profession is on borrowed time. We had a government guy saying that entire swaths of jobs will be displaced before the year ends. And so on.

The interesting thing is, a lot of people believed it then - and a lot of people believe it now. I wonder if we will have the same deja vu in, say, 2029. No, for sure not. It's going to be done and dusted for human thinking by end of this calendar year.

Post reply on HN