Live data from Hacker News

Viewing profile — txyx303

txyx303

HN member
Joined
Thu, Mar 14, 2024, 1:15 AM UTC
HN karma
12
Public activity
8 items

About txyx303

No profile information was provided.

Recent public activity

  1. comment
    Comment #46999892

    Those are scribe lines where you usually would cut out chips which is why it resembles multiple chips. However, they work with TSMC to etch across them.

  2. comment
    Comment #44763678

    feels very low compared to claude/gpt for me

  3. comment
    Comment #41713649

    I don't think they rely on SRAM very much for training. https://cerebras.ai/blog/the-complete-guide-to-scale-out-on-... outlines the memory architecture but it seems like they are …

  4. comment
    Comment #41713604

    Seems like they support training on a bunch of industry standard models. I think most of the customers in the training space tend to be for fine tuning right? The P and T in GPT st…

  5. comment
    Comment #41713490

    afaik they have the current SOTA language models for arabic

  6. comment
    Comment #41713456

    MLPerf brings in exactly zero revenue. If they have sold every chip they can make for the next 2+ years, why would they be diverting resources to MLPerf benchmarking? Artificial an…

  7. comment
    Comment #41371185

    Batched inference will increase your overall throughput, but each user will still be seeing the original throughput number. It's not necessarily a memory vs compute issue in the sa…

  8. comment
    Comment #39699673

    That was more of a WSE-1 problem maybe? They switched to a new compute paradigm (details on their site if you look up "weight streaming") where they basically store the activation …