Viewing profile — kraken12
kraken12
HN member- Joined
- Sat, Oct 12, 2013, 5:25 PM UTC
- HN karma
- 61
- Public activity
- 10 items
- HN profile
- View on Hacker News ↗
About kraken12
Recent public activity
-
comment
Comment #36701942
The 4090 has half the memory bandwidth, so it could not get a 5X gain, it would actually run slower on a memory bound LLM like this.
-
comment
Comment #36700126
Yeah, it is an architectural simulation study, this is what is usually done right at the beginning before resources are allocated to go deep on idea. So in that sense it is imagina…
-
comment
Comment #36700047
Yeah, a preliminary architectural study to sanity check if an idea could potentially pay off.
-
comment
Comment #36700032
Maybe they could do something like AMD's GPU memory stacking, that is good for scaling, and of course they are using many chips not one chip..
-
comment
Comment #36699997
Seems like HN comments have determined that the cost number is not fudged..
-
comment
Comment #36699979
Seems to me the 18 tokens per second from [1] is the throughput and includes the batch size, so I don't think they misread the Deepspeed inference paper. So the chiplet ASIC superc…
-
comment
Comment #36699765
Yep, it's a research paper in comp arch, the initial proof-of-concept study before you go and spend real money on it.
-
comment
Comment #36699736
They are using many chips and taking advantage of the way data flows in LLMs to make it work; so it would be cost-effective unlike Cerebras
- story
- story