Viewing profile — Chamix
Chamix
HN member- Joined
- Tue, Jun 08, 2021, 9:44 PM UTC
- HN karma
- 71
- Public activity
- 42 items
- HN profile
- View on Hacker News ↗
About Chamix
No profile information was provided.
Recent public activity
-
comment
Comment #49177168
Naturally, he has a whole detailed post/trace about it on gwern net! https://gwern.net/twitter An interesting case of echo chamber formation in that its pragmatic to be scared of o…
-
comment
Comment #49177106
Ha, haha, across hundreds of personal discussions I've been involved with on lesswrong/lighthaven, twitter, Wikipedia talk/editing, SF parties etc I think his most distinguishing f…
-
comment
Comment #48216339
I note that (though summarized), this is ~100k tokens. Anyone who routinely works with Codex (or any agentic harness really) can tell you how trivial it is to eat up 100k tokens do…
- comment
-
comment
Comment #47370540
I appreciate the detailed comment! I took the day off and am bored so have a brain dump of a reply - basically I think we are talking past each other on two major points: 1. All th…
-
comment
Comment #47327195
Sorry if that was unclear, I did mean 100Bs as in the next order of magnitude. Even GPT-4 had ~220B active params, though the trend has been towards increased sparsification (lower…
-
comment
Comment #47325911
What do you think labs are doing with the minimum 10TB memory in NvLink 72 systems that were publicly reported to all start coming online in November/December of last year? And why…
-
comment
Comment #47320016
I assure you, the number of people paying to use Qwen3-Max or other similar proprietary endpoints is far less than 1.6 billion.
-
comment
Comment #47319811
I generally agree, back of the napkin math shows H20 cluster of 8gpu * 96gb = 768gb = 768B parameters on FP8 (no NVFP4 on Hopper), which lines up pretty nicely with the sizes of re…
-
comment
Comment #47319659
Try 10s of trillions. These days everyone is running 4-bit at inference (the flagship feature of Blackwell+), with the big flagship models running on recently installed Nvidia 72gp…
-
comment
Comment #46941971
You, know, it sure does add some additional perspective to the original Anthropic marketing materia... ahem, I mean article, to learn that the CCC compiled runtime for SQLite could…
-
comment
Comment #43199215
Indeed, and the difference could in essence be achieved yourself with a different system prompt on 4o. What exactly is 4.5 contributing here in terms of a more nuanced intelligence…
-
comment
Comment #43199076
It's interesting to compare the cost of that original gpt-4 32k(0314) vs gpt-4.5: $60/M input tokens vs $75/M input tokens $120/M output tokens vs $150/M output tokens
- comment
-
comment
Comment #40827835
Forgive me if I'm missing your existing realization (I did a quick check of your HN, reddit, twitter, LW), but I think the big deal with Sohu (wrt Etched) is that they have pivoted…
- comment
-
comment
Comment #40344706
I was thinking about the llm writing tool from Janus.
- comment
-
comment
Comment #39847716
4chan already has a torrent out, of course.
-
comment
Comment #38344885
The little secret is that the training run (meaning, creating the raw autocompleting multimodal token weights) for 5 ran in parallel with 4.
-
comment
Comment #38328538
Luckily Eliezer has written hundreds of approachable essays on the development of his epistemic processes over at lesswrong.com so you too can learn rationality and derive the kill…
-
comment
Comment #38326553
Fair enough, shame "Large Tokenized Models" etc never entered the nomenclature.
-
comment
Comment #38323717
You are conflating Illya's belief in the transformer architecture (with tweaks/compute optimizations) being sufficient for AGI with that of LLMs being sufficient to express human-l…
- comment
-
comment
Comment #36414369
The issue, as pointed above, is primarily bandwidth (at inference), not addressable memory. Put simply, the best bandwidth stack we currently have is on-package HBM -> NVLink, -> M…