Live data from Hacker News

Viewing profile — magic_at_nodai

magic_at_nodai

HN member
Joined
Mon, Jul 01, 2013, 2:47 PM UTC
HN karma
57
Public activity
57 items

About magic_at_nodai

No profile information was provided.

Recent public activity

  1. comment
    Comment #47056503

    yes lmk how i can help. at the minimum i can get you hw and help with PRs etc. firstname at amd.com to reach me.

  2. comment
    Comment #43209512

    Im running ROCm ok on my 9070XT. You can build it from source today if you have a card. rocminfo: **** Agent 2 **** Name: gfx1201 Uuid: GPU-cea119534ea1127a Marketing Name: AMD Rad…

  3. comment
    Comment #42783787

    ROCm on Radeon should work too and the poll above was to seek feedback on what to cards to support next.

  4. comment
    Comment #42783713

    I will provide this feedback to the docs team to clean up. I found it hard when i was making that Poll :D but I looked harder instead of trying to fix the docs. So thank you for th…

  5. comment
    Comment #42783700

    Is this the repo you are referring to https://github.com/amd/go_amd_smi ? Would having a prebuilt version there help you ?

  6. comment
    Comment #42783676

    yes. We are behind on software support for all consumer cards and would love to support all cards. But are looking for guidance / feedback so we can prioritize.

  7. comment
    Comment #42783654

    I have quad w7900s under my desk that work well for workloads on my desktop that translate well to MI300x. There are some perf gaps with FAv2, and FP8 but otherwise I get a seamles…

  8. comment
    Comment #42783594

    We do care about software and acknowledge the gaps and will work hard to make it better. Please let me know any specific issues that are an issue for you and Im happy to push for i…

  9. comment
    Comment #42783560

    PTX does provide a low level machine abstraction. However you still target some version of hardware ( https://arnon.dk/matching-sm-architectures-arch-and-gencode-... ). However a l…

  10. comment
    Comment #42783490

    hey thats me. Happy to help answer anything here and look forward to your constructive feedback to make AMD software better. We got work to do and look forward to it.

  11. comment
    Comment #39568272

    AMD Artificial Intelligence Group (AIG) | Remote / Global AMD Artificial Intelligence Group (AIG) leads AMD AI strategy and drives AI roadmap across client, edge, and cloud. We bui…

  12. comment
    Comment #35115288

    We have it running as part of SHARK (which is built on IREE). https://github.com/nod-ai/SHARK/tree/main/shark/examples/sha...

  13. comment
    Comment #34084016

    Can you give SHARK a try and let us know on our discord? We can try to help. People have been using it on older AMD GPUs back to Polaris arch.

  14. comment
    Comment #34083989

    Here are a list of potential issues https://github.com/AUTOMATIC1111/stable-diffusion-webui/disc... That said we (Nod.ai team) will add support for xformers soon so you can opt in …

  15. comment
    Comment #33733286

    Try SHARK on your AMD GPUs for SD. Follow the setup here: https://github.com/nod-ai/SHARK/tree/main/shark/examples/sha... . It works with Pytorch -> torch-mlir -> MLIR / IREE -> vu…

  16. comment
    Comment #30448056

    unlikely since the interface from ANE is not public and it may change between hardware versions.

  17. comment
    Comment #30437768

    I updated the blog with the reference. Basically it crashes to compile the model with https://github.com/NodLabs/shark-samples/blob/main/examples/... . The coremltools converter is…

  18. comment
    Comment #30435865

    hear your pain and we really want to make it easy (after we make it work). //part of nod.ai / SHARK team.

  19. comment
    Comment #30435856

    So with Tensorcores you use TF32 which is more like FP19-ish and the marketing makes you think you get 8x the performance. But if you want actual FP32 precision you will need somet…

  20. comment
    Comment #30435704

    Yeah the ANE and AMX on cpu are wrapped behind Accelerate Framework and CoreML. So you will have to use CoreML (which wasn't able to compile the latest TF BERT). ANE is also infere…

  21. comment
    Comment #30435648

    Thanks to: LLVM/MLIR --> For the awesome compiler infrastructure IREE --> For the awesome backend to MLIR SHARK/nod.ai --> For adapting IREE for use on various hardware and fine tu…

  22. comment
    Comment #30435616

    This is not part of regular pytorch install. If you can build torch-mlir and SHARK from src you can use it. So hopefully soon we can make pip installable packages but for now the i…

  23. comment
    Comment #30435591

    Here are the matmul sizes for the MiniLM model used for inference: https://github.com/mmperf/mmperf/blob/main/benchmark_sizes/b... These are the matmul sizes for the BERT training …

  24. comment
    Comment #29068683

    nod.ai | Wherever you want in the US | https://nod.ai Come work on A.I Compilers, Runtimes and ML Systems in an _all_ engineer team. You will be working on the forefront on ML Fram…

  25. comment
    Comment #28723233

    nod.ai | Wherever you want in the US | https://nod.ai Come work on A.I Compilers, Runtimes and ML Systems in an _all_ engineer team. No PHBs. We are looking for A.I Compiler Engine…