SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency
infini-ai-lab.github.io
SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency
1–10 of 65 posts
Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency
#2Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency
#3Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency
#4It's almost a mathematical certainty that people who invested in OpenAI will need to reincarnate in multiple universes to ever see that money again but no bother many are probably NVIDIA stock holders to even out the damage.
Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency
#5So this is 8x faster for serving these models than before? Or is this about it being more deterministic? I can't quite tell from reading it.
Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency
#6this is quite worrying for OpenAI as the rate token prices have been plummeting thanks to Meta and its going to have to keep cutting its prices while capex remains flat. whatever Sam says in interviews just think the opposite and the whole picture comes together. It's almost a mathematical certainty that people who invested in OpenAI will need to reincarnate in multiple universes to ever see that money again but no b…
Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency
#7this is quite worrying for OpenAI as the rate token prices have been plummeting thanks to Meta and its going to have to keep cutting its prices while capex remains flat. whatever Sam says in interviews just think the opposite and the whole picture comes together. It's almost a mathematical certainty that people who invested in OpenAI will need to reincarnate in multiple universes to ever see that money again but no b…
Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency
#8Will this work, or do I need a Tesla P40 or two?
Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency
#9Re: SEQUOIA: Exact Llama2-70B on an RTX4090 with half-second per-token latency
#10I'm looking at buying 2 X RTX 3060s to run LLama 70b for my new PC I just purchased. Will this work, or do I need a Tesla P40 or two?
You will also be getting about 720GB/s of memory bandwidth with 2x3060; instead of 1TB/s with the 4090; so expect lower performance.