Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

61–70 of 712 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#61

Imagine the size of chip needed to 'etch' something like Qwen 3.6 27B in size.

Not too dissimilar to the first HC1 (6nm 815mm² 53B Transistors embedding an 8b LLM):

> Our second model, still based on Taalas’ first-generation silicon platform (HC1), will be a mid-sized reasoning LLM

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#62

Earlier quoted context omitted.

I know it's a relatively tiny model, but damn, is that thing fast. It also mostly passes the "schlong" test https://pastes.io/YcxSi8Fp

I didn't realize there was a SchlongBench™ (but of course there is). What's it test? (asking seriously)

There isn't SchlongBench(TM) yet, it's a specific question I've been asking of differently sized models as a randomly chosen gauge of how much less commonly used knowledge is perma-baked into it. In this case a question about a specific yiddish origin slang term. Small/bad models don't know it's from middle high german or Yiddish and get its origin and meaning totally wrong (or it runs into model censorship related to slang related to the male anatomy).

It's also a question I have found will cause models that don't know what it is to go off quickly in a direction of hallucination trying to explain it, so the hallucination is evident very quickly starting from the first ever prompt issued with 0 context fill. Example: I had a model write four detailed supposedly-accurate sounding, grammatically correct paragraphs saying its origin is from AAVE (African American Vernacular English), which it most certainly is not

You could do the same by picking any topic that is very rarely discussed in conversation, some esoteric and narrow piece of knowledge and asking the model about it.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#63

Earlier quoted context omitted.

Which model? Or how many active parameters?

Llama 3.1 8B model

So this demo is around 90 times faster than typical speeds for the same model at openrouter, and around 30 times faster than the absolute fastest option available (Groq).

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#64
post #52

Can anyone comment on the economics and likely turnaround times of this process, when it’s more mature? Would it be realistic for a frontier lab to deploy this or would the turnaround time mean the model is always too out of date? Assuming the weights and architecture are eventually stable, how much cheaper would this end up being?

There are always uses for outdated models.

Claude Code is still using haiku 4.5 from ages ago for explore subagents for instance. Not to mention production uses like customer service that only need to be "good enough"

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#65
post #52

Can anyone comment on the economics and likely turnaround times of this process, when it’s more mature? Would it be realistic for a frontier lab to deploy this or would the turnaround time mean the model is always too out of date? Assuming the weights and architecture are eventually stable, how much cheaper would this end up being?

I mean even if it take a few months, it'll still be out of date. But there was a hypothetical when it came up in Feb, would you want Qwen 3.5 at like 10k tokens per second.

At the time people were no doubt saying yes but now 3.8 is out, is that still desirable?

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#66
post #58

Imagine the size of chip needed to 'etch' something like Qwen 3.6 27B in size.

If someone has that sort of knowledge; how big a chip would be required? Is it possible?

Well, given the data above, roughly a 220b transistors chip for the HC1 tech.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#67

Honestly, this is starting to make more and more sense. SOTA models are starting to converge to certain architecture and capabilities. I wouldn’t be surprised we end up with a base model ASIC + “fine tune” card where it’s a physical LoRA style adapter.

Imagine a multi-modal model with 1000's of tokens per second. Realtime inference for a host of applications. This is a BIG deal and will change the landscape in unfathomable ways. The https://chatjimmy.ai demo was impressive. Once models settle down this makes sense. Imagine a cartridge with a physical model on it. You purchase a cartridge and stick it in your computer/phone/server. Want to upgrade? By a new 'cartrid…

i'm looking forward to Qwen3.8 27B launch to see how much models have peaked at a given size.

it might already be time to start burning the best small models onto hardware since it's possible they can't get much better at many tasks like knowledge recall due to the inherent information density limits for models at a given size.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#68
post #60

I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.

Considering the rate of model development and rail hopping, seems like baking models into silicon is speed-running obsolescence.

Not sure. You can fix the transistors but leave the connections between them open for flexibility, so you only need to change the manufacturing process for the upper masks for every new model.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#69

Imagine the size of chip needed to 'etch' something like Qwen 3.6 27B in size.

Interesting thought, because it's a yield question. How tolerant are models today to a few broken weights.

If tolerant, they could churn out many cheaper chips, some perhaps with slight abnormal tendencies ;)

Post reply on HN