Instant AI Response
chatjimmy.ai
Instant AI Response
1–10 of 15 posts
Re: Instant AI Response
#2What model and hardware powers this?
Is this a Google T5 based model?
Re: Instant AI Response
#3This is a demo of Taalas inference ASIC hardware. Prior discussion @ https://news.ycombinator.com/item?id=47086181
Re: Instant AI Response
#4[deleted]
Re: Instant AI Response
#5Re: Instant AI Response
#6What model and hardware powers this? Is this a Google T5 based model?
3bit hard-wired Llama 3.1 8B ( https://taalas.com/the-path-to-ubiquitous-ai/ )
Re: Instant AI Response
#7If this is possible, why not all online AI engines work like this?
Re: Instant AI Response
#8I love seeing optimised SLM inference. Is there a current use-case for this? Edge CNNs make sense to me but not edge SLMs (yet).
Re: Instant AI Response
#9imagine a model like opus 4.6 at that speed, that would be insane
Re: Instant AI Response
#10If this is possible, why not all online AI engines work like this?
This is an specific model (Llama 3.1 8B) baked in hardware form. You can only use this model but get "low" power consumption and crazy speed.
If you want to run a different model you need new hardware for that new model.