Live data from Hacker News

The path to ubiquitous AI (17k tokens/sec)

taalas.com

461–470 of 471 posts

Re: The path to ubiquitous AI (17k tokens/sec)

#461

Earlier quoted context omitted.

> they'd fail on any novel problem not in their training data Yes, and that's exactly what they do. No, none of the problems you gave to the LLM while toying around with them are in any way novel.

None of my codebases are in their training data, yet they routinely contribute to them in meaningful ways. They write code that I'm happy with that improves the codebases I work in. Do you not consider that novel problem solving?

Correct, you are not doing any novel problem solving.

Re: The path to ubiquitous AI (17k tokens/sec)

#462

Earlier quoted context omitted.

I see, that's very cool, that's the context I was missing, thanks a lot for explaining.

I don't mean to be rude, but did you read the article before commenting?

I'm commenting on the link to their demo, not on the article.

Re: The path to ubiquitous AI (17k tokens/sec)

#463
post #442

The NextPlatform article hints at their approach: “We have got this scheme for the mask ROM recall fabric – the hard-wired part – where we can store four bits away and do the multiply related to it – everything – with a SINGLE TRANSISTOR. So the density is basically insane. And this is not nuclear physics – it is fully digital. It is just a clever trick that we don’t want to broadcast. But once you hardwire everythin…

TIL your salary

(Kiddin’, my silly way to say thanks for a deeply technical look, helps me understand the kind of knowledge work that might be useful n years from now!)

Re: The path to ubiquitous AI (17k tokens/sec)

#464
post #406
post #400

Earlier quoted context omitted.

Yea I mean this is the first publishable draft of a startup cooking on this. I'm confident there are at least 1-2 OOMs of improvement to come here in terms of the (intelligence : wattage) ratio. I really thought we were going to need to see a couple of dramatic OOM-improvement changes to the model composition / software layer, in order to get models of Opus 3.7's capability running on our laptops. This release tells…

The way I imagine it in 2-4 years we're going to be hit with a triple glut of better architecture, massive oversupply of hardware and potentially one or two hardware efforts like this really taking off. It's pretty crazy we're already 4 years in and outside of very niche / low availability solutions, it's still either GPU or bust

That's interesting! How do you see "oversupply of hardware" playing out?

Is it because we stop doing ~2024-style, large-scale training (marginal returns aren't worth it)? Or because supply way outpaces the training+inference demand?

AFAIU if the trend lines /S-curves keep chugging along as they are, we won't hit hardware oversupply for a long, long time without some sort of AI training winter.

Re: The path to ubiquitous AI (17k tokens/sec)

#465

There's an old idea of adaptive media. Imagine a video drama that's composed of a graph of clips, like an old "choose your own adventure" book ("Do you X? If yes, goto page 45"). With gaze tracking, one can "hmm, the viewer is more focused on character A than B... so we'll give clips and subplots with more A". Now, when reading, the eye moves in little jumps - saccades. They last 10's of ms, the eye is blind during t…

Generative TikTok for words

Hmm... TikTok has apparently long had "text enhanced with background" genres, and TIL, text posts since 2023. So text is ok. But non-independent items? For generative storytelling, "here is a next paragraph for the story", swipe left/right might work? Want to avoid "I don't much like this new paragraph, but I'm afraid to lose it and be stuck with something worse". Swipe left/right and up for continue? Swipe down to revisit old choices? Maybe present new text bolded, appended to old text, for context. Or a "next page of a picture book" idiom. A text field for direct creative or editorial intervention - speech to text. Maybe a side channel input for "story and background should now be soporific". Generative bedtime stories, but incrementally collaboratively created... Thanks for the brainstorming prompt.

Re: The path to ubiquitous AI (17k tokens/sec)

#466
post #390

Earlier quoted context omitted.

Yeah, just cause Cisco had a huge market lead on telecom in the late '90s, it doesn't mean they kept it. (And people nowadays: "Who's Cisco?")

They did mostly keep it though.

Sure, but it's taken their stock price about 20 years to recover.

Re: The path to ubiquitous AI (17k tokens/sec)

#467

If I could have one of these cards in my own computer do you think it would be possible to replace claude code? 1. Assume It's running a better model, even a dedicated coding model. High scoring but obviously not opus 4.5 2. Instead of the standard send-receive paradigm we set up a pipeline of agents, each of whom parses the output of the previous. At 17k/tps running locally, you could effectively spin up tasks like…

Models can't improve themselves with their own (model) input, they need to be grounded in truth and reality.

But at one point the model is sufficiently large enough to accomplish any task a human could specify. For software development, I think we're pretty much at that point with the latest Anthropic/Google/OpenAI models. We have no idea where the direction of token pricing is going to go in the future, but the consensus seems to be that it will only get more expensive. If Taalas can offer the same functionality that we have with frontier models today at a 1/10 of the cost and 10x the speed then they're going to take over a large part of the market.

Re: The path to ubiquitous AI (17k tokens/sec)

#468
post #160

Earlier quoted context omitted.

Models don’t get old as fast as they used to. A lot of the improvements seem to go into making the models more efficient, or the infrastructure around the models. If newer models mainly compete on efficiency it means you can run older models for longer on more efficient hardware while staying competitive. If power costs are significantly lower, they can pay for themselves by the time they are outdated. It also means…

From my own experience, models are at the tipping point for being useful at prototypes in software, and those are very large frontier models not feasible to get down on wafers unless someone does something smart. I really don't like the hallucination rate for most models but it is improving, so that is still far in the future. What I could see though, is if the whole unit they made would be power efficient enough to…

> From my own experience, models are at the tipping point for being useful at prototypes in software

You must not have much experience using the new frontier models then. A lot of large tech companies are replacing their SDLC with agentic workflows. The tooling and frameworks are still ramping up, but the models have no problem producing production ready software given proper specifications.

Re: The path to ubiquitous AI (17k tokens/sec)

#469

I'm curious how much of "hardcoding" is in the chip? Can it have parts that don't need changing much and "offload" the rest into some sort of high-speed/bandwidth interconnect? Will we reach a state where we have chips on which models can be "flashed" like CPU firmware? Or eventually will we reach a state where none of these tricks will be needed because like run-of-the mill Intel/AMD commodity CPUs, we will have ful…

These chips are large by fab standards and even with state of the art processes we likely won't see any kind of integration on consumer tech any time soon, but I imagine they will absolutely see instant demand if they can deliver on what they laid out in the post.

Re: The path to ubiquitous AI (17k tokens/sec)

#470
post #433

Earlier quoted context omitted.

There is nothing new here. This has been demonstrated several times by previous researchers: https://arxiv.org/abs/2511.06174 https://arxiv.org/abs/2401.03868 For a real world use case, you would need an FPGA with terabytes of RAM. Perhaps it'll be a Off chip HBM. But for s large models, even that won't be enough. Then you would need to figure out NV-link like interconnect for these FPGAs. And we are back to square o…

This is new. You are citing FPGA prototypes. Those papers do not demonstrate the same class of scaling or hardware integration that Taalas is advocating. For one, the FPGA solutions typically use fixed multipliers (or lookup tables), the ASIC solution has more freedom to optimize routing for 4 bit multiplication.

I understand that what Taalas is claiming. I was trying to actually describe that model on a hardware is some not something new Or unthought of The natural progression of FPGA is ASIC. Taalas process is more expensive And not really worth it because once you burn a model on the silicon, the silicon can only serve that model. speed improvement alone is not enough for the cost you will incur in the long run. GPU's are still general purpose, FPGA's are atleast reusable but wont have the same speed. But this alone cannot be a long term business. Turning a model to hardware in two months is too long. Models already take quite a long time to train. Anyone going down this strategy would leave wide open field to their competitors. Deployment planning of existing models already so complicated.
Post reply on HN