Earlier quoted context omitted.
"seems like baking models into silicon is speed-running obsolescence" Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.
Well, 50TB ROM Taalas HC1 style would be apparently a 400000b transistor system through a chip sized 2.5 meters on the side... :)
AMD acquires Taalas to boost inference performance by etching models in silicon
291–300 of 712 posts
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#292It must be a “super model”. What will be if new model released? New chips?
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#293Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#294Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.
Perfect for consumers. You buy it and then you need to buy a new one in a couple of years. If they can make them affordable they'll sell like hotcakes.
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#295The demo: https://chatjimmy.ai/
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#296Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#297Given the fast churn of the models, how does it work out? Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out. Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.
If they begin etching Fable into silicon now and release it 2-3 years later, i can see the market for it
Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#298I'm surprised neither OpenAI nor Anthropic made this move first. The Chinese open weight models are pulling ahead and commoditizing their value proposition. Baking models onto silicon would've been the next logical move to get a moat. Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
Just to see how fast it is try chatjimmy.ai
"Yes, the Wang Corporation, the company that originally developed and marketed the Wang 2200 computer, still exists as a rebranded company under the name PPL (Precision Pencil and Label), but it has undergone significant changes and challenges over the years.
Here's a brief overview of what happened:
Founding and Growth: The Wang Corporation was founded by An Wang in 1969."
In fact, Wang labs was founded in 1951. PPL seems to be a made up entity. But it did generate those "facts" in 0.033 seconds. If people value speed over accuracy then I can write an LLM that is 100x faster than chatjimmy.ai and make big bucks by responding one of N canned responses to any question.Re: AMD acquires Taalas to boost inference performance by etching models in silicon
#299Earlier quoted context omitted.
This is only true for people who are solely focused on performance. There is absolutely a market for acceptable performance combined with predictability.
True, but predictability cuts both ways. We're all used to having to constantly update our browsers and phones to keep up with the security arms race. If a frozen model can't be updated, it will predictably remain vulnerable to any "exploits" or idiosyncratic quirks that people discover over time. Let's say, as somebody suggested in another comment, that you buy 100,000 of these chips and deploy them to run fast-food…
Does it though? Isn't that what CPUs are, very fast-not-so-clever computing brain surrounded by layers that protect it?