Earlier quoted context omitted.
I had the same reaction but then I showed it to my partner. She completely didn't get it, in her words "how can it be thinking of a good answer when it's that quick?" I tried to explain but I fear were probably going to be adding artificial sleeps to these things to convince the masses it's doing something clever.
It's not thinking. Not in the way she probably meant. It can "think" that fast the same way a calculator can "think" that fast (kind of). Because it's not human and not "thinking", it's a mathematical algorithm
AMD acquires Taalas to boost inference performance by etching models in silicon
441–450 of 729 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#442Earlier quoted context omitted.
Baking the base models on to ROM makes a lot of economic sense. Less so for consumers though, because it'd mean the phone is out of date in 3 months when a better model comes along.
It's a perfect reason to get consumers to buy a new phone every year again! They got bored of the camera.
Now you sell the same phone with higher price tag.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#443Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#444Earlier quoted context omitted.
The Taalas chips are not physically small. And part of their secret (if you look at the design) is just locating a bunch of memory soldered on the edges ( I belive higher amounts of SRAM ? )
Baking the base models on to ROM makes a lot of economic sense. SRAM for the KV cache & fine-tunes, not so much. Sure you’d get incredible speeds but it’s not scalable from a die-size or cost perspective. Rather base model on ROM + KV cache on DRAM is much more scalable. Also this would work great for edge devices that have a 2-5 year lifecycle.
Add a bunch of chips together, and you get to a server that can run a 800B model, very fast and probably significantly cheaper than others.
[1]https://www.eetimes.com/taalas-specializes-to-extremes-for-e...
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#445Earlier quoted context omitted.
What would possibly tell you that?
Expensive as fuck to make chips, only makes sense if you believe whatever model you're creating a chip out of will not become completely irrelevant in 5-10 years.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#446Earlier quoted context omitted.
Depends on how much it costs the consumer. If I could buy a "cartridge" of Kimi K3 for 300 bucks I 100% would buy that shit asap. Even if it's "no good" after lets say 4 months still would be worth it IMO.
That's definitely super-enthousiast territory. Paying 80 bucks a month for AI is more than 99.99% of people would be willing to do
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#447Earlier quoted context omitted.
What would possibly tell you that?
Expensive as fuck to make chips, only makes sense if you believe whatever model you're creating a chip out of will not become completely irrelevant in 5-10 years.
There are probably limited applications but not zero.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#448Earlier quoted context omitted.
Plug it in, and it's a old prototype with Gemma 5 weights baked onboard. Dammit, fucked by Craigslist again!
Back in the kazaa and limewire days, you'd sometimes try to get a movie / episode from a series, wait hours / days for it to download, and when it was done you had a ~50/50 chance to actually watch what you wanted or an old german porn movie :/
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#449Thinking that five or six years from now, Fable-level intelligence could be provided at 100x the current speed... makes me feel lost. I cannot imagine what the future will look like.
It makes me dizzy. I have no idea what is going to happen within even a year from now, can barely even imagine it. I'm trying to get the most out of it by redlining my AI subscriptions. Hopefully I'll manage to start a business in my niche. I don't even know if my niche will exist in the future.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#450Earlier quoted context omitted.
What, even if it means you can run models without relying on the currently backlogged DRAM production?
The size of model we're talking about running doesn't need much if any dram.
I could see it being feasible to get a Qwen-3.6-27b type of model done on something like this. Qwen-3.6-27b at 18tok/s would be a game changer.