Viewing profile — txyx303
txyx303
HN member- Joined
- Thu, Mar 14, 2024, 1:15 AM UTC
- HN karma
- 12
- Public activity
- 8 items
- HN profile
- View on Hacker News ↗
About txyx303
No profile information was provided.
Recent public activity
-
comment
Comment #46999892
Those are scribe lines where you usually would cut out chips which is why it resembles multiple chips. However, they work with TSMC to etch across them.
-
comment
Comment #44763678
feels very low compared to claude/gpt for me
-
comment
Comment #41713649
I don't think they rely on SRAM very much for training. https://cerebras.ai/blog/the-complete-guide-to-scale-out-on-... outlines the memory architecture but it seems like they are …
-
comment
Comment #41713604
Seems like they support training on a bunch of industry standard models. I think most of the customers in the training space tend to be for fine tuning right? The P and T in GPT st…
-
comment
Comment #41713490
afaik they have the current SOTA language models for arabic
-
comment
Comment #41713456
MLPerf brings in exactly zero revenue. If they have sold every chip they can make for the next 2+ years, why would they be diverting resources to MLPerf benchmarking? Artificial an…
-
comment
Comment #41371185
Batched inference will increase your overall throughput, but each user will still be seeing the original throughput number. It's not necessarily a memory vs compute issue in the sa…
-
comment
Comment #39699673
That was more of a WSE-1 problem maybe? They switched to a new compute paradigm (details on their site if you look up "weight streaming") where they basically store the activation …