Viewing profile — dhruvdh
dhruvdh
HN member- Joined
- Fri, Apr 26, 2019, 6:36 PM UTC
- HN karma
- 557
- Public activity
- 153 items
- HN profile
- View on Hacker News ↗
About dhruvdh
No profile information was provided.
Recent public activity
-
comment
Comment #46646472
Try `uvx pocket-tts serve`
- story
-
comment
Comment #43463796
El Capitan can also do FP8. HPC requires double precision generally but people are trying to make low precision work.
-
comment
Comment #43205571
To be fair, you can buy ~3 of these for the price Nvidia charges for 24GB/32GB models.
- comment
-
comment
Comment #42773748
To add, AMD only makes _parts_ of an MI300X server. It's like asking a tire manufacturer to give you a car for free.
-
comment
Comment #42516218
I wish more people would just try to do things just like this and blog about their failures. > The published version of a proof is always condensed. And even if you take all the ma…
-
comment
Comment #42491771
Disappointed that there wasn’t anything on inference performance in the article at all. That’s what the major customers have announced they use it for.
-
comment
Comment #42491739
Which algorithm you pick for what shape of matrices is different and not straightforward to figure out. AMD currently wants you to “tune” ops and likely search for the right algori…
- story
-
comment
Comment #42056312
> despite them being fabless That's not how it works. You need to pump money into fabs to get them working, and Intel doesn't have money. If AMD had fabs to light up their money, t…
-
comment
Comment #42055249
> Performance per watt was better for Intel No, not its not even close. AMD is miles ahead. This is a Phoronix review for Turin (current generation): https://www.phoronix.com/revie…
-
comment
Comment #42022873
Is AMD behind hyperscaler in-house efforts? Outside of Google I don't think so.
-
comment
Comment #41865253
Oh, maybe also change the title? I flagged it because of the title/url not matching.
-
comment
Comment #41809422
I don't think having a common ancestry for the ISA means much, or even having the same ISA. Anyway, I don't understand what you want from me or are arguing about. They were trying …
-
comment
Comment #41809225
Those are Vega, not CDNA. It wouldn't surprise me if those are rebranded consumer chips, though I haven't checked.
-
comment
Comment #41809099
And yet Meta is using MI300X exclusively for all live inference on Llama 405B. Clearly there are workloads AMD wins at, and just going Nvidia by default for everything without cons…
-
comment
Comment #41809066
You know AMD primarily sells CPUs right? For datacenter GPUs, they're going from ~500M-750M in 2023 full year (can't find proper numbers), to 4.5B+ full year 2024. In GPUs, it's al…
-
comment
Comment #41724901
Batching is how you get ~350 tokens/sec on Qwen 14b on vLLM (7900XTX). By running 15 requests at once. Also, there is a Dockerfile.rocm at the root of vLLM's repo. How is it a pain…
-
comment
Comment #41723376
Why would you use this over vLLM?
-
comment
Comment #40504914
What's the point of the 8000 LOC limit? Has anyone worked in a project with a LOC limit? Why was the limit in place?
-
comment
Comment #40435718
The MacBook NPU is 3x slower than the 45 TOPS threshold required for Copilot+ PC branding.
-
comment
Comment #40434310
Yeah, and any updates to the model that make Recall attractive, cannot go back in time and reanalyze the past to be more useful. At least they will learn a lot from this, in 6-12 m…
-
comment
Comment #40434200
The NPU runs this Silica model at 1.5 watts. MacBooks cannot even drive multiple monitors in this price range.
-
comment
Comment #40318346
The FPGA being used is I believe one of the lowest speced SKUs. AWS instance prices are more of a supply/demand/availability thing, it would be more interesting to compare from a t…