Viewing profile — somnial
somnial
HN member- Joined
- Thu, Oct 30, 2025, 10:18 AM UTC
- HN karma
- 10
- Public activity
- 16 items
- HN profile
- View on Hacker News ↗
About somnial
No profile information was provided.
Recent public activity
- story
-
comment
Comment #49171387
throughput scales superlinearly with number of GPUs when networked well and deployed with wideEP, so 1x won't compare. also it would be interesting to figure from the DSpark paper …
- story
- story
- story
-
comment
Comment #48699802
this is a blog post from a company that hosts open weights LLMs ( https://www.doubleword.ai/ ). I think its possible it might have been tongue in cheek
-
comment
Comment #48491672
https://fergusfinn.com/blog/economics-of-speculative-decodin... good point tho - plus for Deepseek the shared expert increases the overlap slightly
-
comment
Comment #48443249
true, but no reason the predictor model couldn't use linear attention (i.e. mamba, GDN etc) to predict KV caches
- story
- story
- story
- story
- story
- story
- story
- story