Live data from Hacker News

The path to ubiquitous AI (17k tokens/sec)

taalas.com

51–60 of 471 posts

Re: The path to ubiquitous AI (17k tokens/sec)

#51
post #43

Edit: it seems like this is likely one chip and not 10. I assumed 8B 16bit quant with 4K or more context. This made me think that they must have chained multiple chips together since N6 850mm2 chip would only yield 3GB of SRAM max. Instead, they seem to have etched llama 8B q3 with 1k context instead which would indeed fit the chip size. This requires 10 chips for an 8 billion q3 param model. 2.4kW. 10 reticle sized…

> What is a task that is extremely high value, only require a small model intelligence, require tremendous speed, is ok to run on a cloud due to power requirements, AND will be used for years without change since the model is etched into silicon? Video game NPCs?

Doesn’t pass the high value and require tremendous speed tests.

Re: The path to ubiquitous AI (17k tokens/sec)

#52
I think the thing that makes 8b sized models interesting is the ability to train unique custom domain knowledge intelligence and this is the opposite of that. Like if you could deploy any 8b sized model on it and be this fast that would be super interesting, but being stuck with llama3 8b isn't that interesting.

Re: The path to ubiquitous AI (17k tokens/sec)

#53

Edit: it seems like this is likely one chip and not 10. I assumed 8B 16bit quant with 4K or more context. This made me think that they must have chained multiple chips together since N6 850mm2 chip would only yield 3GB of SRAM max. Instead, they seem to have etched llama 8B q3 with 1k context instead which would indeed fit the chip size. This requires 10 chips for an 8 billion q3 param model. 2.4kW. 10 reticle sized…

ceo

No one would never give such a weak model that much power over a company.

Re: The path to ubiquitous AI (17k tokens/sec)

#54
post #13

This would be killer for exploring simultaneous thinking paths and council-style decision taking. Even with Qwen3-Coder-Next 80B if you could achieve a 10x speed, I'd buy one of those today. Can't wait to see if this is still possible with larger models than 8B.

It uses 10 chips for 8B model. It’d need 80 chips for an 80b model. Each chip is the size of an H100. So 80 H100 to run at this speed. Can’t change the model after you manufacture the chips since it’s etched into silicon.

Do we know that it needs 10 chips to run the model? Or are the servers for the API and chatbot just specced with 10 boards to distribute user load?

Re: The path to ubiquitous AI (17k tokens/sec)

#55
post #33

Earlier quoted context omitted.

Where are those numbers from? It's not immediately clear to me that you can distribute one model across chips with this design. > Model is etched onto the silicon chip. So can’t change anything about the model after the chip has been designed and manufactured. Subtle detail here: the fastest turnaround that one could reasonably expect on that process is about six months. This might eventually be useful, but at the mo…

> The first generation HC1 chip is implemented in the 6 nanometer N6 process from TSMC. Each HC1 chip has 53 billion transistors on the package, most of it very likely for ROM and SRAM memory. The HC1 card burns about 200 watts, says Bajic, and a two-socket X86 server with ten HC1 cards in it runs 2,500 watts. https://www.nextplatform.com/2026/02/19/taalas-etches-ai-mod...

So it lights money on fire extra fast, AI focused VCs are going to really love it then!!

Re: The path to ubiquitous AI (17k tokens/sec)

#56
post #46

Earlier quoted context omitted.

Don’t forget that the 8B model requires 10 of said chips to run. And it’s a 3bit quant. So 3GB ram requirement. If they run 8B using native 16bit quant, it will use 60 H100 sized chips.

> Don’t forget that the 8B model requires 10 of said chips to run. Are you sure about that? If true it would definitely make it look a lot less interesting.

Their 2.4 kW is for 10 chips it seems based on the next platform article.

I assume they need all 10 chips for their 8B q3 model. Otherwise, they would have said so or they would have put a more impressive model as the demo.

https://www.nextplatform.com/2026/02/19/taalas-etches-ai-mod...

Re: The path to ubiquitous AI (17k tokens/sec)

#57
Strange that they apparently raised $169M (really?) and the website looks like this. Don't get me wrong: Plain HTML would do if "perfect", or you would expect something heavily designed. But script-kiddie vibe coded seems off.

The idea is good though and could work.

Re: The path to ubiquitous AI (17k tokens/sec)

#58
The article doesn't say anything about the price (it will be expensive), but it doesn't look like something that the average developer would purchase.

An LLM's effective lifespan is a few months (ie the amount of time it is considered top-tier), it wouldn't make sense for a user to purchase something that would be superseded in a couple of months.

An LLM hosting service however, where it would operate 24/7, would be able to make up for the investment.

Re: The path to ubiquitous AI (17k tokens/sec)

#59
post #21

This is not a general purpose chip but specialized for high speed, low latency inference with small context. But it is potentially a lot cheaper than Nvidia for those purposes. Tech summary: - 15k tok/sec on 8B dense 3bit quant (llama 3.1) - limited KV cache - 880mm^2 die, TSMC 6nm, 53B transistors - presumably 200W per chip - 20x cheaper to produce - 10x less energy per token for inference - max context size: flexib…

Were we go towards really smart roboters. It is interesting what kind of diferent model chips they can produce.
Post reply on HN