Live data from Hacker News

Viewing profile — eden-u4

eden-u4

HN member
Joined
Mon, Nov 18, 2024, 11:53 PM UTC
HN karma
23
Public activity
28 items

About eden-u4

No profile information was provided.

Recent public activity

  1. comment
    Comment #47713666

    I don't understand the RMS table, shouldn't it be non commutative? Eg "example 0 vs 1"'s RMS != "example 1 vs 0"'s RMS? Which doesn't seem the case for the checkpoints I checked.

  2. comment
    Comment #47664968

    There are not enough samples in that book to generate new "infinite" data.

  3. comment
    Comment #47664817

    Guardrailing is usually done with a smaller model (< 1b) to filter out simple "not aligned prompt" and not waste compute.

  4. comment
    Comment #47389199

    But what if your rubber duck is actually steering your thought process (since you may not have a consolidated one)? In this way I think the AI as editor is far better than a rubber…

  5. comment
    Comment #47150959

    This is basically a diffusion model: start from a random seed, use a generative process to transform it into something

  6. story
  7. comment
    Comment #46411590

    mine too, but none was such a dick. also, anything related to school (particularly at a young age), is not viewed as something to boast of (at least in my experience in italy, serb…

  8. comment
    Comment #46125720

    Also, they have all the infra to actually use all that tpus advantage (as well as actual researchers, contrariwise to OpenAI)

  9. comment
    Comment #45963081

    what type of experiments did you run in less than a week to be so dismissing? (seriously curious)

  10. story
  11. comment
    Comment #45860148

    Not OP, but I guess based on your comment: > But the extent and the way in which Zig specifically puts it to use -- which includes, but is not limited to, how it is used to replace…

  12. comment
    Comment #45832620

    wow, thanks for this long explanation.

  13. comment
    Comment #45071929

    so they are basically using a similar idea to that of a stirling engine in thermoelectric generator or they use a different mechanism to produce energy?

  14. comment
    Comment #44491810

    No open model/weights?

  15. comment
    Comment #44471354

    I don't have much experience with ROCm for large trainings, but NVIDIA is still shit with driver+cuda version+other things. The only simplification is due to ubuntu and other distr…

  16. comment
    Comment #44002248

    why don't you ask the model about the shrinked system prompt and the original system prompt? in this way you can infer whether the same relevant informations are "stored" in the hi…

  17. comment
    Comment #43829764

    I dunno, these reasoning models seems kinda "dumb" because they try to bootstrap itself via reasoning, even though a simple direct answer might not exist (for example key informati…

  18. story
  19. comment
    Comment #42830981

    When they say "in 20-50 years" it is implied that it might not exist ever. See nuclear fusion or AGI.

  20. comment
    Comment #42830962

    I think the issue with RL is that, in order for a model to perform well in a task, you have to make it stubborn. In the same way a student that thinks outside the scope of the task…

  21. comment
    Comment #42771122

    because non-tech friends don't have coding problems, which is the only problem LLM are ok at.

  22. comment
    Comment #42653843

    anyone knows which program they used to record the demo?

  23. comment
    Comment #42548793

    meta is large enough (and diversified enough), when OpenAI valuation goes to pennies (because LLM won't solve any real problem, nor they'll remain relevant in the long term and ope…

  24. comment
    Comment #42520922

    this project only uses kaggle metadata and abstract from arxiv. Moreover it is "focused" on only 5-6 categories in the arxiv. Therefore, the costs are marginal. Plus you could use …

  25. comment
    Comment #42468499

    if you try hard enough you can always ~latinize~ ASCIIfy a language