Live data from Hacker News

Viewing profile — Chamix

Chamix

HN member
Joined
Tue, Jun 08, 2021, 9:44 PM UTC
HN karma
71
Public activity
42 items

About Chamix

No profile information was provided.

Recent public activity

  1. comment
    Comment #49177168

    Naturally, he has a whole detailed post/trace about it on gwern net! https://gwern.net/twitter An interesting case of echo chamber formation in that its pragmatic to be scared of o…

  2. comment
    Comment #49177106

    Ha, haha, across hundreds of personal discussions I've been involved with on lesswrong/lighthaven, twitter, Wikipedia talk/editing, SF parties etc I think his most distinguishing f…

  3. comment
    Comment #48216339

    I note that (though summarized), this is ~100k tokens. Anyone who routinely works with Codex (or any agentic harness really) can tell you how trivial it is to eat up 100k tokens do…

  4. comment
  5. comment
    Comment #47370540

    I appreciate the detailed comment! I took the day off and am bored so have a brain dump of a reply - basically I think we are talking past each other on two major points: 1. All th…

  6. comment
    Comment #47327195

    Sorry if that was unclear, I did mean 100Bs as in the next order of magnitude. Even GPT-4 had ~220B active params, though the trend has been towards increased sparsification (lower…

  7. comment
    Comment #47325911

    What do you think labs are doing with the minimum 10TB memory in NvLink 72 systems that were publicly reported to all start coming online in November/December of last year? And why…

  8. comment
    Comment #47320016

    I assure you, the number of people paying to use Qwen3-Max or other similar proprietary endpoints is far less than 1.6 billion.

  9. comment
    Comment #47319811

    I generally agree, back of the napkin math shows H20 cluster of 8gpu * 96gb = 768gb = 768B parameters on FP8 (no NVFP4 on Hopper), which lines up pretty nicely with the sizes of re…

  10. comment
    Comment #47319659

    Try 10s of trillions. These days everyone is running 4-bit at inference (the flagship feature of Blackwell+), with the big flagship models running on recently installed Nvidia 72gp…

  11. comment
    Comment #46941971

    You, know, it sure does add some additional perspective to the original Anthropic marketing materia... ahem, I mean article, to learn that the CCC compiled runtime for SQLite could…

  12. comment
    Comment #43199215

    Indeed, and the difference could in essence be achieved yourself with a different system prompt on 4o. What exactly is 4.5 contributing here in terms of a more nuanced intelligence…

  13. comment
    Comment #43199076

    It's interesting to compare the cost of that original gpt-4 32k(0314) vs gpt-4.5: $60/M input tokens vs $75/M input tokens $120/M output tokens vs $150/M output tokens

  14. comment
  15. comment
    Comment #40827835

    Forgive me if I'm missing your existing realization (I did a quick check of your HN, reddit, twitter, LW), but I think the big deal with Sohu (wrt Etched) is that they have pivoted…

  16. comment
  17. comment
    Comment #40344706

    I was thinking about the llm writing tool from Janus.

  18. comment
  19. comment
    Comment #39847716

    4chan already has a torrent out, of course.

  20. comment
    Comment #38344885

    The little secret is that the training run (meaning, creating the raw autocompleting multimodal token weights) for 5 ran in parallel with 4.

  21. comment
    Comment #38328538

    Luckily Eliezer has written hundreds of approachable essays on the development of his epistemic processes over at lesswrong.com so you too can learn rationality and derive the kill…

  22. comment
    Comment #38326553

    Fair enough, shame "Large Tokenized Models" etc never entered the nomenclature.

  23. comment
    Comment #38323717

    You are conflating Illya's belief in the transformer architecture (with tweaks/compute optimizations) being sufficient for AGI with that of LLMs being sufficient to express human-l…

  24. comment
  25. comment
    Comment #36414369

    The issue, as pointed above, is primarily bandwidth (at inference), not addressable memory. Put simply, the best bandwidth stack we currently have is on-package HBM -> NVLink, -> M…