Earlier quoted context omitted.
If apple offers these in a data center that’s open access it’s game over for NVDA
Based on which fantasy premise?
and somewhat competitive on cost, but that wont be the main selling point, just availability
21–30 of 47 posts
Earlier quoted context omitted.
If apple offers these in a data center that’s open access it’s game over for NVDA
Based on which fantasy premise?
and somewhat competitive on cost, but that wont be the main selling point, just availability
With "borrow" in scare quotes, do you intend for him to be aware of his generosity?
With "borrow" in scare quotes, do you intend for him to be aware of his generosity?
Probably just because in the full phrasing, ``i would like to "borrow" his gpu``, omitting the scare quotes would paint the picture of their friend unplugging/unsoldering their GPU and lending it to the author.
For llama3 just ask him to install ollama and serve the model. Ollama has auto memory management and will free the model when not used, and whenever you make a call to the API (do let your friend know before you do this) ollama will reload the model back to memory again. Not sure whether there are anything similar for SD though.
> i do not want to remote into his machine > tailscale/zerotier Same thing isn't it? In any case it wouldn't be hard for you to just have an account on his machine, tailscale being perhaps the simplest setup. SSH in, cook his laptop at your leisure.
You make http requests to the shared server, those get proxied via the ssh tunnel to his machine, and the client on his machine could make the determination when/whether to run the workload.
An off-topic question, are Apple's M-series chips any good at current AI/ML work? How does it compare with dedicated GPUs?
I have an M3 chip in my laptop, it has more memory than my 4090 but it's still way slower when inferencing. So as long as the model fits in memory, Nvidia GPUs are going to be way faster just because they have more/faster compute cores. Of course, if the model fits in memory on your M chip and doesn't in your Nvidia chip, the M chip wins by default. However, I would say, if you load a 70B model in your M chip, while…
Earlier quoted context omitted.
Based on which fantasy premise?
The premise where these are readily available for mass purchase, have a hardware and software stack that already works reliably, and have a lower energy footprint than other offerings and somewhat competitive on cost, but that wont be the main selling point, just availability
Earlier quoted context omitted.
Based on which fantasy premise?
The premise where these are readily available for mass purchase, have a hardware and software stack that already works reliably, and have a lower energy footprint than other offerings and somewhat competitive on cost, but that wont be the main selling point, just availability