If this is possible, why not all online AI engines work like this?
This is an specific model (Llama 3.1 8B) baked in hardware form. You can only use this model but get "low" power consumption and crazy speed. If you want to run a different model you need new hardware for that new model.
Instant AI Response
11–15 of 15 posts
Re: Instant AI Response
#12Re: Instant AI Response
#13What model and hardware powers this? Is this a Google T5 based model?
3bit hard-wired Llama 3.1 8B ( https://taalas.com/the-path-to-ubiquitous-ai/ )
Re: Instant AI Response
#14Earlier quoted context omitted.
This is an specific model (Llama 3.1 8B) baked in hardware form. You can only use this model but get "low" power consumption and crazy speed. If you want to run a different model you need new hardware for that new model.
It is really a crazy speed. 15k tokens/second.
Generated in 0,074s • 15 754 tok/s
Re: Instant AI Response
#15If this is possible, why not all online AI engines work like this?
This is an specific model (Llama 3.1 8B) baked in hardware form. You can only use this model but get "low" power consumption and crazy speed. If you want to run a different model you need new hardware for that new model.