Performance per dollar is getting faster and cheaper
141–150 of 151 posts
Re: Performance per dollar is getting faster and cheaper
#142I think we should make it illegal to not specify the quantization in the headline for these types of posts.
I do think the phrasing is weird, though. The performance may be improving, but it isn't the thing getting "faster" (e.g. responses to queries might get "faster"). And the dollars aren't getting cheaper; the performance is. "Performance per dollar" is a rate; it is not getting faster or cheaper. It should just say "Performance per dollar is increasing".
Re: Performance per dollar is getting faster and cheaper
#143I think we should make it illegal to not specify the quantization in the headline for these types of posts.
I don't know what you mean by "quantization". I guess you're asking to explain how they measure "performance", but that kind of thing often won't neatly fit in a title. I do think the phrasing is weird, though. The performance may be improving, but it isn't the thing getting "faster" (e.g. responses to queries might get "faster"). And the dollars aren't getting cheaper; the performance is. "Performance per dollar" is…
The old headline was something like "We served GLM5.2 on AMD MI355X at 2626 tok/s/node". They think it's bad to advertise performance numbers like that without specifying how much you quantized the model (Since you can 'cheat' more performance by quantizing more, and not mentioning the reduced output quality)
Re: Performance per dollar is getting faster and cheaper
#144Earlier quoted context omitted.
What does it mean it's better at nvfp4 training? What's different between training and inference to make this true?
I'm also puzzled by that statement. The issue with training is (as I understand it) one of precision and the associated numerical stability. You need enough bits in order for backprop to function correctly. Of course there are techniques such as quantization aware training but I don't understand why a datatype would work for inference but not for that. You can also abandon backprop entirely but that comes with a whol…
Re: Performance per dollar is getting faster and cheaper
#145While cool, quantization to FP4 is practically never lossless in actual use. A lot of providers are advertising high TPS on Kimi and GLM, but the models are functionally lobotomized and no longer close to frontier quality. Would love to see this not be true.
Doesn't Nvidia with their NVFP4 claim that it's lossless? I haven't tested enough models Nvidia has converted to NVFP4 besides GLM 5.2 but it seemed fine to me. My own luck has been hit or miss with it.
Re: Performance per dollar is getting faster and cheaper
#146This is very interesting and yet not at the same time. This looks to be optimized for single-stream LLM traffic which is not viable to serve in a production setting. It's only interesting to hobbyists that want to run the model locally. It's genuinely neat that AI can find the right optimization pathways in an AMD inference server to unlock this but at the same token (pun-intended) this is a classic case of benchmark…
hi yes it’s not optimized for single stream it’s optimized for total node throughput
Re: Performance per dollar is getting faster and cheaper
#147Earlier quoted context omitted.
The fact that people have lived and worked near data centres for decades and didn't even know what the term meant - let alone be adversely impacted by them - probably indicates they're broadly an non issue. All of a sudden out of nowhere, AI and data centres got intermingled by the media and now people seem to have big issues with them.
Because the dynamics have shifted enormously inside the rack. 10 years ago, I was running 4 CPU servers with 48 cores and 128GB of RAM in 2U enclosures with a maximum power consumption of 500W or so. I was able to stick ~20 of them in a 42U rack, totaling 10kW. A data center full of these can be cooled with CRACs and hot/cold aisles without much problem. This is still too much for a bog-standard server colocation ope…
That kind of energy density is scary.
[1] https://blog.se.com/datacenter/2025/10/16/the-1-mw-ai-it-rac...
Re: Performance per dollar is getting faster and cheaper
#148Re: Performance per dollar is getting faster and cheaper
#149Earlier quoted context omitted.
You say they have a large impact, but having lived somewhere with some of the largest data centers- they very much don't. At least not more so then any other structure that paves over greenery. love to debate actual discission points. pull up "datacenter dfw" on google maps for mine.
The people having glass literally break from the vibrations would probably disagree with your opinion https://youtu.be/_bP80DEAbuo?is=sg09k66iutKFIFSo Yet here we are, discussing "data center" as if they're standardized and of similar (nose) isolation. There are no meaningful regulations in building them, and they can be incredibly polluting. So your experience with a potentially well isolated one is sadly not the no…
So you want to use the internet but dont want the data center in your backyard?
Re: Performance per dollar is getting faster and cheaper
#150While cool, quantization to FP4 is practically never lossless in actual use. A lot of providers are advertising high TPS on Kimi and GLM, but the models are functionally lobotomized and no longer close to frontier quality. Would love to see this not be true.
Doesn't Nvidia with their NVFP4 claim that it's lossless? I haven't tested enough models Nvidia has converted to NVFP4 besides GLM 5.2 but it seemed fine to me. My own luck has been hit or miss with it.
https://www.reddit.com/r/LocalLLaMA/comments/1twz9ur/cyankiw...