Live data from Hacker News

AMD acquires Taalas to boost inference performance by etching models in silicon

theregister.com

351–360 of 712 posts

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#352
post #334
post #180

Earlier quoted context omitted.

Wait, is it even thinking? Or is it an instant model?

It’s not reasoning, the hardware demo uses a 3.-something generation Llama 8B. But it’s proven they can automate this (they didn’t etch eight billion weights by hand after all, obviously), so now the interesting question is whether they can scale it to more recent aka bigger models. After all, there’s already very useful models even for productivity at 27 or 35B.

My concern is that reasoning could involve some sequential steps that instant models don't.

Not sure if modern models "think" only by outputting blocks, or there is a more complex mechanism at play.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#353
post #241

Earlier quoted context omitted.

I know it's a relatively tiny model, but damn, is that thing fast. It also mostly passes the "schlong" test https://pastes.io/YcxSi8Fp

It failed on my usual test. But it failed really fast: "A farmer has a wolf, a goat, and a cabbage. The wolf is imaginary and doesn't exist. He wants to cross the river, but the boat is only big enough to hold him and one of them. The farmer can't leave the wolf and the goat together, because the wolf will eat the goat. Similarly, he can't leave the goat and the cabbage together, because the goat will eat the cabbage…

This farmer needs a tote.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#354

This is neat but IMO a little crazy. Something I personally haven’t seen much of, in all the discussions of model benchmarks and AI breakthroughs, is a distinction between “peak performance” and “reliable performance”. The “peak performance” of frontier models is very high: they’re solving open math problems, analyzing large codebases, etc. But my subjective impression is that “reliable performance” is mid at best: o…

I think you're underestimating both their reliability for standard problems and the usefulness of that level of reliability.

This is a good point. Opus does some silly shenanigans sometimes but then catches it later. It’s still an order of magnitude faster at getting to a working system than I am, for ones I don’t know.

It’s really a dream for setting up a homelab

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#355

Earlier quoted context omitted.

Depends on how much it costs the consumer. If I could buy a "cartridge" of Kimi K3 for 300 bucks I 100% would buy that shit asap. Even if it's "no good" after lets say 4 months still would be worth it IMO.

That's definitely super-enthousiast territory. Paying 80 bucks a month for AI is more than 99.99% of people would be willing to do

That's because the super-enthusiast will upgrade in 4 months when a better model is released. The casual user would keep it for years. A year of claude at the lowest plan is almost $300

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#356
post #202

Earlier quoted context omitted.

There's a benchmark for this and a lot of models get negative scores because they're so unreliable: https://artificialanalysis.ai/evaluations/omniscience

I wanted some examples they actually experienced. Because I use these things daily and haven’t seen a hallucination in a long long time.

Search a terminal with Claude Code for things like, “I got it wrong twice. I should look up the documentation instead of guessing.”

Does it about once a day, that I notice.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#357
post #278

Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.

Cerebras already runs large models like Kimi 2.6 or GLM at like 30x speed. 100 times is next year, not six years. You can actually test it out on their website, just imagine 3 x faster and maybe 15% smarter.

Cerebras is literally the entire wafer, so it can't get bigger. So where is the jump from 30x to 100x coming from? Node improvements only yield like 10-20% gains these days...

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#358
post #185

Earlier quoted context omitted.

I think this would make sense for consumer hardware, not for AI companies. AI companies constantly update/change stuff, new models come out, new requirements, etc. But if you ship an "ai-powered" dishwasher, it can come with the chip built-in to do computer vision and precisely target each spot, and will be sold as-is with no updates.

You don’t need this chip to do that. Computer vision has used machine learning for decades. The task you’re describing is pretty rudimentary and an off the shelf model with a control system would do it way cheaper.

Think of a HomePod. 99% (and likely much more) of what people are asking is super simple.

Re: AMD acquires Taalas to boost inference performance by etching models in silicon

#360
post #170

Earlier quoted context omitted.

obsolescence is the whole point. apple gets to sell a new phone very 6-12 months because of it. i have written about this: "For device makers Packaging models with laptops and smartphones will let application access near free, low latency inference and potentially offer users a better experience with the option of preserving data on-device. This is viable under the condition that tasks that do require larger expert m…

No, the point is inference speed and power.

you don't understand what I wrote.
Post reply on HN