Live data from Hacker News

The path to ubiquitous AI (17k tokens/sec)

taalas.com

421–430 of 471 posts

Re: The path to ubiquitous AI (17k tokens/sec)

#421

Earlier quoted context omitted.

If it's so easy to do custom silicon for any model (they say only 2 months), why didn't they demo one of the newer DeepSeek models instead? Using a 2-year model is so bad. I'm not buying it.

they explain it in the article: this is the first iteration, so they wanted to start with something simple, ie, this is a tech demo.

Ok then I look forward to seeing DeepSeek running instantly at the end of April.

Re: The path to ubiquitous AI (17k tokens/sec)

#422
post #418
post #375

Earlier quoted context omitted.

Did you see the part in my original post where it said "Not unexpected for an 8k model"?

Oh I saw it, you still have a fundamentally flawed comprehension of LLMs. The size of the model does not factor as tiny models can use Internet to fetch factual information. But you think they are accurate repositories of knowledge, even though it's physically impossible unless lossless infinite compression algorithms exist (they don't, can't and won't).

I think you're overestimating your ability to assess what others think or comprehend.

Re: The path to ubiquitous AI (17k tokens/sec)

#425

If I could have one of these cards in my own computer do you think it would be possible to replace claude code? 1. Assume It's running a better model, even a dedicated coding model. High scoring but obviously not opus 4.5 2. Instead of the standard send-receive paradigm we set up a pipeline of agents, each of whom parses the output of the previous. At 17k/tps running locally, you could effectively spin up tasks like…

It's 2.5kW so it likely won't sit in your computer (quite beyond what a desktop could provide in power alone to a single card, let alone cool). It's 8.5cm^2 which is a beast of a single die. Basically logistically it's going to need to be in a data centre. It's ideal for small context high throughput. Perhaps parsing huge text piles like if you had the entire Epstein files as text. I think Claude code benefits from l…

> It's 2.5kW so it likely won't sit in your computer (quite beyond what a desktop could provide in power alone to a single card, let alone cool). It's 8.5cm^2 which is a beast of a single die.

I wonder how you cool a 3x3cm die that outputs 2.5 kW of heat. In the article they mention that the traditional setup requires water cooling, but surely this does as well, right?

Re: The path to ubiquitous AI (17k tokens/sec)

#426
> Taalas’ silicon Llama achieves 17K tokens/sec per user, nearly 10X faster than the current state of the art, while costing 20X less to build, and consuming 10X less power.

Am I reading this right: 10x faster and 10x less power, ie. 100x more power efficient?

Re: The path to ubiquitous AI (17k tokens/sec)

#427

Earlier quoted context omitted.

they explain it in the article: this is the first iteration, so they wanted to start with something simple, ie, this is a tech demo.

Ok then I look forward to seeing DeepSeek running instantly at the end of April.

Why so negative lol. The speed and very reduced power use of this thing are nothing to be sneezed at. I mean, hardware accelerated LLMs are a huge step forward. But yeah, this is a proof of concept, basically. I wouldn't be surprised if the size factor and the power use go down even more, and that we'll start seeing stuff like this in all kinds of hardware. It's an enabler.

Re: The path to ubiquitous AI (17k tokens/sec)

#428
post #294

Earlier quoted context omitted.

And then it's slow again to finally find a correct answer...

It won't find the correct answer. Garbage in, garbage out.

How about if you run this loop (one year from now) on this kind of hardware but with something like Claude/Kimi K2. How about that? Because that's where it'll go.

Re: The path to ubiquitous AI (17k tokens/sec)

#429

Earlier quoted context omitted.

A related argument I raised a few days back on HN: What's the moat with with these giant data-centers that are being built with 100's of billions of dollars on nvidia chips? If such chips can be built so easily, and offer this insane level of performance at 10x efficiency, then one thing is 100% sure: more such startups are coming... and with that, an entire new ecosystem.

You'd still need those giant data centers for training new frontier models. These Taalas chips, if they work, seem to do the job of inference well, but training will still require general purpose GPU compute

Yeah but you need even bigger factories to fabricate those inference chips, so what is the point?

Re: The path to ubiquitous AI (17k tokens/sec)

#430
post #276

Earlier quoted context omitted.

Yes there are some fascinating emergent properties at play, but when they fail it's blatantly obvious that there's no actual intelligence nor understanding. They are very cool and very useful tools, I use them on a daily basis now and the way I can just paste a vague screenshot with some vague text and they get it and give a useful response blows my mind every time. But it's very clear that it's all just smoke and mi…

When humans fail a task, it’s obvious there is no actual intelligence nor understanding. Intelligence is not as cool as you think it is.

It can still be cool- but maybe it's just not as rare.
Post reply on HN