Live data from Hacker News

Hacked Nvidia 4090 GPU driver to enable P2P

github.com

321–330 of 365 posts

Re: Hacked Nvidia 4090 GPU driver to enable P2P

#321
post #134

Earlier quoted context omitted.

Any time I open guys steam half of it is some sort of politics

You can blame chat for that lol

He should ban chat and focus on development. Leave the political talk to the kids in their respective Discord servers.

Re: Hacked Nvidia 4090 GPU driver to enable P2P

#322

Earlier quoted context omitted.

tinygrad supports uneven splits. There's no fundamental reason for 4 or 8, and work should almost fully parallelize on any number of GPUs with good software. We chose 6 because we have 128 PCIe lanes, aka 8 16x ports. We use 1 for NVMe and 1 for networking, leaving 6 for GPUs to connect them in full fabric. If we used 4 GPUs, we'd be wasting PCIe, and if we used 8 there would be no room for external connectivity asid…

That is very interesting if tinygrad can support it! Every other library I've seen had the limitation on dividing the heads, so I'd (perhaps incorrectly) assumed that it's a general problem for inference.

There are some interesting hacks you can do like replicating the K/V weights by some factor which allows them to be evenly divisible by whatever number of gpus you have. Obviously there is a memory cost there, but it does work.

Re: Hacked Nvidia 4090 GPU driver to enable P2P

#323
post #314

Earlier quoted context omitted.

Did you at least front run the market and stocked up of 4090ies before this release? Also gamers are probably not too happy about these developments :D

4090's have consistently been around 2000 dollars. I don't think there's many gamers who would be affected by price fluctuations of the 4090 or even the 4080.

This is out of touch; they were mad before and they will be mad again. Lots of people spend a huge chunk of their modest disposable income on high end gaming gear, and the only upside of these issues for them is that eventually, YEARS down the line, capacity/supply issues MIGHT calm down in a way that yields some benefits.

They're going to realize soon enough that they've basically just been told that the extremely shitty problem they thought they'd moved beyond is back with a vengeance and the next generation of gaming cards has the potential to make the past few rounds of scalping shit-shows look tame.

Re: Hacked Nvidia 4090 GPU driver to enable P2P

#324

Earlier quoted context omitted.

Yes, and I am sure that when people do a google search for "Good arguments in favor of X", that they are also sometimes convinced to be more in favor of X. Perhaps they would be even more convinced by the google search than if a person argued with them about it. That is still much different from "The AI mind controls people, hacks the nukes, and ends the world". Its that second part that is the the fantasy land situa…

If the only evidence for AI doom you will accept is actual AI doom, you are asking for evidence that by definition will be too late. "Show me the AI mindcontrolling people!" AI mindcontrolling people is what we're trying to avoid seeing. The trick is, in the world in which AI doom is in the future, what would you expect to see now that is different from the world in which AI doom is not in the future?

> If the only evidence for AI doom you will accept is actual AI doom

No actually. This is another mistake that the AI doomers make. They pretend like a demand for evidence means that the world has to end first.

Instead, what would be perfectly good evidence, would be evidence of significant incremental harm that requires regulation on its own, independent of any doom argument.

In between "the world literally ends by magic diamond nanobots and mind controlling AI" and "where we are today" would be many many many situations of incrementally escalating and measurable harm that we would see in real life, decades before the world ending magic happens.

We can just treat this like any other technology, and regulate it when it causes real world harm. Because before the world ends by magic, there would be significant real world harm that is similar to any other problem in the world that we handle perfectly well.

Its funny because you committing the exact mistake that I was criticizing in my original post, where you did the absolutely massive jump and hand waved it away.

> what would you expect to see now that is different from the world in which AI doom is not in the future?

What I would expect is for the people who claim to care about AI doom to actually be trying to measure real world harm.

Ironically, I think the people who are coming up with increasingly thin excuses as for why they don't have to find evidence are increasing the likelyhood of such AI doom much more than anyone else because they are abandoning the most effective method of actually convincing the world of the real world damage that AI could cause.

Re: Hacked Nvidia 4090 GPU driver to enable P2P

#325

Earlier quoted context omitted.

This matches my own take. I've tuned into a few of his streams and watched VODs on YouTube. I am consistently underwhelmed by his actual engineering abilities. He is that particular kind of engineer that constantly shits on other peoples code or on the general state of programming yet his actual code is often horrendous. He will literally call someone out for some code in Tinygrad that he has trouble with and then he…

link your github. want to see your raw intellectual power

It's over 9000.

Re: Hacked Nvidia 4090 GPU driver to enable P2P

#326

So assuming you utilized this with (4) x 4090s is there a theoretical comparative to performance vs the A6000 / other professional lines?

It depends on what you do with it and how much bandwidth it needs between the cards. For LLM inference with tensor parallelism (usually limited by VRAM read bandwidth, but little exchange needed) 2x 4090 will massively outperform a single A6000. For training, not so much.

Re: Hacked Nvidia 4090 GPU driver to enable P2P

#327

Earlier quoted context omitted.

> surely it's better to have distributed AGI than to have AGI only in the hands of the elites. The argument of doing so is the same as Nuclear Non-Proliferation - because of its great abuse potential, giving the technology to everyone only causes random bombings of cities instead of creating a system with checks and balances. I do not necessarily agree with it, but I found the reasoning is not groundless.

But the reason for nuclear non-proliferation is to hold onto power. Abuse potential is a great excuse, but it applies to everyone . Current nuclear states have demonstrated that they are willing to indirectly abuse them (you can't invade Russia, but Russia has no problem invading you as long as you aren't backed up by nukes).

Both can be true at the same time.

The world's superpowers enforce nuclear non-proliferation mainly because it allows them to keep unfair political and military advantages to themselves. At the same time, one cannot deny that centralized weapon ownership made the use of such weapons more controllable: These nuclear states are powerful enough to establish a somewhat responsible chain of command to avoid their unreasonable or accidental uses, and so far these attempts are still successful. Also, due to the fact that they are "too big to fail", they were forced to hire experts to make detailed analysis on the consequences of nuclear wars, and the resulted MAD doctrine discouraged them from starting such wars.

On the other hand, if the same nuclear technologies are available to everyone, the chance of an unreasonable or accidental nuclear war will be higher. If even resourceful superpowers can barely keep these nuclear weapons under safe political and technical control (as shown by multiple incidents and near-misses during the Cold War [0]), surely a less resourceful state or military in possession of equally destructive weapons will have even more difficulties on controlling their uses.

At least this is how the argument goes (so far, I personally take no position).

Of course, I clearly realized that centralized control is not infallible. Months ago, in a previous thread on OpenAI's refusal on publishing technical details of GPT-4, most people believed that they were using it as an excuse to maintain a monopolistic control. Instead, I argued that perhaps OpenAI truly values the problem of safety right now - but acting responsibly right now is not an indication that they will still act responsibly in the future. There's no guarantee that the safety considerations will eventually be overridden in favor of financial gains.

[0] https://en.wikipedia.org/wiki/Command_and_Control_(book)

Re: Hacked Nvidia 4090 GPU driver to enable P2P

#328

Earlier quoted context omitted.

If the only evidence for AI doom you will accept is actual AI doom, you are asking for evidence that by definition will be too late. "Show me the AI mindcontrolling people!" AI mindcontrolling people is what we're trying to avoid seeing. The trick is, in the world in which AI doom is in the future, what would you expect to see now that is different from the world in which AI doom is not in the future?

> If the only evidence for AI doom you will accept is actual AI doom No actually. This is another mistake that the AI doomers make. They pretend like a demand for evidence means that the world has to end first. Instead, what would be perfectly good evidence, would be evidence of significant incremental harm that requires regulation on its own, independent of any doom argument. In between "the world literally ends by…

Well, at least if you see escalating measurable harm you'll come around, I'm happy about that. You won't necessarily get the escalating harm even if AI doom is real though, so you should try to discover if it is real even in worlds where hard takeoff is a thing.

> What I would expect is for the people who claim to care about AI doom to actually be trying to measure real world harm.

Why bother? If escalating harm is a thing, everyone will notice. We don't need to bolster that, because ordinary society has it handled.

Re: Hacked Nvidia 4090 GPU driver to enable P2P

#330

Earlier quoted context omitted.

> but there's nothing "suboptimal non-default options" about it If "bypassing the official driver to invoke the underlying hardware feature directly through source code modification (and incompatibilities must be carefully worked around by turning off IOMMU and large BAR, since the feature was never officially supported)" does not count as "suboptimal non-default options", then I don't know what counts as "suboptimal…

> bypassing the official driverto The driver is not bypasses. This is a patch to the official open-source kernel-driver where the feature is added, which is how all upstream Linux driver development is done. > to invoke the underlying hardware feature directly Accessing hardware features directly is pretty much the sole job of a driver, and the only thing "bypassed" is some abstractions internal to the driver. Just m…

> The driver is not bypasses. This is a patch to the official open-source kernel-driver where the feature is added, which is how all upstream Linux driver development is done. [source code modification] is a weird way to describe software engineering. Making the code available for further development is kind of the whole point of open source.

My previous comment was written with an unspoken assumption: Hardware drivers tend to be very different from other forms of software. For ordinary free-and-open-source software, the source code availability largely guarantees community control. However, the same often does not apply to drivers. Even with source code availability, they're often written by vendors using NDAed information and in-house expertise about the underlying hardware design. As a result, drivers remain under a vendor's tight control. Even with access to 100% source code, it's often still difficult to do meaningful development due to missing documentation to explain "why" instead of "what", the driver can be full of magic numbers and unexplained functionalities, without any description other than a few helpers functions and macros. This is not just a hypothetical scenario, this situation is encountered by OpenBSD developers on a daily basis. In a OpenBSD presentation, the speaker said the study of Linux code is a form of "reverse-engineering from source code".

Geohot didn't find the workaround by reading hardware documentation, instead, it was found by making educated guesses based on the existing source code, and by watching what happens when you send the commands to hardware to invoke a feature unexposed by the HAL. Thus, it was found by reverse-engineering (in a wider sense). And I call it a driver bypass, in the sense that it bypasses the original design decisions made by Nvidia's developers.

> [turning off IOMMU] is not a P2PDMA problem, and just a result of them not also adding the necessary IOMMU boilerplate, which would be added if the patch was done properly to be upstreamed.

Good point, I stand corrected.

I'll consider stop calling geohot's hack "a bypass" and accepting your characterization of "driver development" if it really gets upstreamed to Linux - which usually requires maintainer review, and Nvidia's maintainer is likely to reject the patch.

> [large BAR] is an expected and "optimal" system requirement.

I meant "turning off (IOMMU && large BAR)". Disabling large BAR in order to use PCIe P2P is a suboptimal configuration.

Post reply on HN