Live data from Hacker News

Instant AI Response

chatjimmy.ai

11–15 of 15 posts

Re: Instant AI Response

#11
post #7

If this is possible, why not all online AI engines work like this?

This is an specific model (Llama 3.1 8B) baked in hardware form. You can only use this model but get "low" power consumption and crazy speed. If you want to run a different model you need new hardware for that new model.

It is really a crazy speed. 15k tokens/second.

Re: Instant AI Response

#13
post #6
post #2

What model and hardware powers this? Is this a Google T5 based model?

3bit hard-wired Llama 3.1 8B ( https://taalas.com/the-path-to-ubiquitous-ai/ )

3bit is a bit ridiculous. From that page I am unclear if the current model is 3 or 4bit. If it’s 4bit… well, NVIDIA showed that a well organized model can perform almost as well as 8bit.

Re: Instant AI Response

#14
post #11

Earlier quoted context omitted.

This is an specific model (Llama 3.1 8B) baked in hardware form. You can only use this model but get "low" power consumption and crazy speed. If you want to run a different model you need new hardware for that new model.

It is really a crazy speed. 15k tokens/second.

I have tried it again. This is the future of chat UI, imho.

Generated in 0,074s • 15 754 tok/s

Re: Instant AI Response

#15
post #7

If this is possible, why not all online AI engines work like this?

This is an specific model (Llama 3.1 8B) baked in hardware form. You can only use this model but get "low" power consumption and crazy speed. If you want to run a different model you need new hardware for that new model.

Do we understand how to scale up the hardware to the point it can run a frontier model? Because this is insane. It will be a game changer for agent systems making 10-100+ calls.
Post reply on HN