Live data from Hacker News

Viewing profile — kraken12

kraken12

HN member
Joined
Sat, Oct 12, 2013, 5:25 PM UTC
HN karma
61
Public activity
10 items

About kraken12

Interested in chip design and semiconductor industry.

Recent public activity

  1. comment
    Comment #36701942

    The 4090 has half the memory bandwidth, so it could not get a 5X gain, it would actually run slower on a memory bound LLM like this.

  2. comment
    Comment #36700126

    Yeah, it is an architectural simulation study, this is what is usually done right at the beginning before resources are allocated to go deep on idea. So in that sense it is imagina…

  3. comment
    Comment #36700047

    Yeah, a preliminary architectural study to sanity check if an idea could potentially pay off.

  4. comment
    Comment #36700032

    Maybe they could do something like AMD's GPU memory stacking, that is good for scaling, and of course they are using many chips not one chip..

  5. comment
    Comment #36699997

    Seems like HN comments have determined that the cost number is not fudged..

  6. comment
    Comment #36699979

    Seems to me the 18 tokens per second from [1] is the throughput and includes the batch size, so I don't think they misread the Deepspeed inference paper. So the chiplet ASIC superc…

  7. comment
    Comment #36699765

    Yep, it's a research paper in comp arch, the initial proof-of-concept study before you go and spend real money on it.

  8. comment
    Comment #36699736

    They are using many chips and taking advantage of the way data flows in LLMs to make it work; so it would be cost-effective unlike Cerebras

  9. story
  10. story