Live data from Hacker News

Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't

vettedconsumer.com

11–20 of 62 posts

Re: Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't

#11
post #10

[flagged]

Maybe like… focus on the actual content instead of perceived writing patterns? Crazy I know.

This. I don't particularly like the LLM writing style, but we've read a ton of very poorly written texts over the year with no complaints. LLMs are average writers with an annoying style, but not bad writers. If the content is good, I don't care if a LLM wrote it.

Re: Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't

#12
post #11
post #10

Earlier quoted context omitted.

Maybe like… focus on the actual content instead of perceived writing patterns? Crazy I know.

This. I don't particularly like the LLM writing style, but we've read a ton of very poorly written texts over the year with no complaints. LLMs are average writers with an annoying style, but not bad writers. If the content is good, I don't care if a LLM wrote it.

It’s a flag that no care went into it. Web surfing involves lots of little decisions following cues of “is this worth my time”

Re: Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't

#13
post #10

[flagged]

Maybe like… focus on the actual content instead of perceived writing patterns? Crazy I know.

The writing style is this staccato LLM-like style that is difficult to read as it has zero flow and meaningless sentences.

Like what does the second sentence even mean? Is it even a sentence? "The roofline math, the prompt-processing catch, the NPU red herring, and the owner-measured speeds."

Re: Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't

#14

Let's also ensure the SSD doesn't age prematurely.

I was under the impression that when you're streaming the weights from disk because the full model won't fit in memory, that it is solely reading from the SSD, not writing, so it wouldn't be causing wear on your SSD.

It is and it doesn't. You only get into disk writes if the system starts paging out to disk.

Re: Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't

#15
post #5

Can't really run it as well, though. My "mini PC" is an M4 Max with 128GB of unified memory and the memory bandwidth is still sorely lacking for inference (although it's far better than any non-unified consumer architecture!).

To be fair it's "only" half the throughput of a 4090 and a third of an RTX 6000. Significant but not an order of magnitude.

An old ada Rtx 6000 maybe. A Blackwell RTX Pro 6000 is an order of magnitude faster and has 96gb.

Re: Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't

#16
post #5

Can't really run it as well, though. My "mini PC" is an M4 Max with 128GB of unified memory and the memory bandwidth is still sorely lacking for inference (although it's far better than any non-unified consumer architecture!).

To be fair it's "only" half the throughput of a 4090 and a third of an RTX 6000. Significant but not an order of magnitude.

Those are the ratios for memory bandwidth, but the GPUs have a much higher ratio for compute, and that affects prefill rate / TTFT, right?

Re: Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't

#17
The current “big GPU” has 96gb of memory, but that’s not a “consumer GPU” apparently, while a $5000 Spark is a “consumer PC” I guess. In any case you’re probably better off running a large open weights model on the cloud.

Re: Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't

#19
post #11

Earlier quoted context omitted.

This. I don't particularly like the LLM writing style, but we've read a ton of very poorly written texts over the year with no complaints. LLMs are average writers with an annoying style, but not bad writers. If the content is good, I don't care if a LLM wrote it.

It’s a flag that no care went into it. Web surfing involves lots of little decisions following cues of “is this worth my time”

It's worse than no care... I read an article recently that was written with no care and it was refreshing.

Putting an LLM on it means you care to make it look nice, but not enough to actually do it. Why bother?

Re: Unified Memory, Explained: Why Mini PCs Can Run 70B Models a Big GPU Can't

#20
post #10

Earlier quoted context omitted.

Maybe like… focus on the actual content instead of perceived writing patterns? Crazy I know.

The writing style is this staccato LLM-like style that is difficult to read as it has zero flow and meaningless sentences. Like what does the second sentence even mean? Is it even a sentence? "The roofline math, the prompt-processing catch, the NPU red herring, and the owner-measured speeds."

Also so sick of the "it's not x, it's y", like stop fucking telling me what it's not! just get to the point!

` The mini PC's slowness is not a driver problem or a weak chip. It is arithmetic on the bandwidth number. `

Post reply on HN