Earlier quoted context omitted.
It uses 10 chips for 8B model. It’d need 80 chips for an 80b model. Each chip is the size of an H100. So 80 H100 to run at this speed. Can’t change the model after you manufacture the chips since it’s etched into silicon.
Do we know that it needs 10 chips to run the model? Or are the servers for the API and chatbot just specced with 10 boards to distribute user load?
The path to ubiquitous AI (17k tokens/sec)
61–70 of 471 posts
Re: The path to ubiquitous AI (17k tokens/sec)
#62I think the thing that makes 8b sized models interesting is the ability to train unique custom domain knowledge intelligence and this is the opposite of that. Like if you could deploy any 8b sized model on it and be this fast that would be super interesting, but being stuck with llama3 8b isn't that interesting.
Model intelligence is, in many ways, a function of model size. A small model tuned for a given domain is still crippled by being small.
Some things don't benefit from general intelligence much. Sometimes a dumb narrow specialist really is all you need for your tasks. But building that small specialized model isn't easy or cheap.
Engineering isn't free, models tend to grow obsolete as the price/capability frontier advances, and AI specialists are less of a commodity than AI inference is. I'm inclined to bet against approaches like this on a principle.
Re: The path to ubiquitous AI (17k tokens/sec)
#63I tried the chatbot. jarring to see a large response come back instantly at over 15k tok/sec I'll take one with a frontier model please, for my local coding and home ai needs..
Reminds me of that solution to Fermi's paradox, that we don't detect signals from extraterrestrial civilizations because they run on a different clock speed.
Re: The path to ubiquitous AI (17k tokens/sec)
#64try here, I hate llms but this is crazy fast. https://chatjimmy.ai/
Re: The path to ubiquitous AI (17k tokens/sec)
#65This is not a general purpose chip but specialized for high speed, low latency inference with small context. But it is potentially a lot cheaper than Nvidia for those purposes. Tech summary: - 15k tok/sec on 8B dense 3bit quant (llama 3.1) - limited KV cache - 880mm^2 die, TSMC 6nm, 53B transistors - presumably 200W per chip - 20x cheaper to produce - 10x less energy per token for inference - max context size: flexib…
Were we go towards really smart roboters. It is interesting what kind of diferent model chips they can produce.
Re: The path to ubiquitous AI (17k tokens/sec)
#66…for a privileged minority, yes, and to the detriment of billions of people whose names the history books conveniently forget. AI, like past technological revolutions, is a force multiplier for both productivity and exploitation.
Re: The path to ubiquitous AI (17k tokens/sec)
#67It could give a boost to the industry of electron microscopy analysis as the frontier model creators could be interested in extracting the weights of their competitors.
The high speed of model evolution has interesting consequences on how often batches and masks are cycled. Probably we'll see some pressure on chip manufacturers to create masks more quickly, which can lead to faster hardware cycles. Probably with some compromises, i.e. all of the util stuff around the chip would be static, only the weights part would change. They might in fact pre-make masks that only have the weights missing, for even faster iteration speed.
Re: The path to ubiquitous AI (17k tokens/sec)
#68This is like microcontrollers, but for AI? Awesome! I want one for my electric guitar; and please add an AI TTS module...
Re: The path to ubiquitous AI (17k tokens/sec)
#69Re: The path to ubiquitous AI (17k tokens/sec)
#70Earlier quoted context omitted.
Absolute insanity to see a coherent text block that takes at least 2 minutes to read generated in a fraction of a second. Crazy stuff...
Yes, but the quality of the output leaves to be desired. I just asked about some sports history and got a mix of correct information and totally made up nonsense. Not unexpected for an 8k model, but raises the question of what the use case is for such small models.