Live data from Hacker News

Viewing profile — mxwsn

mxwsn

HN member
Joined
Fri, Sep 15, 2017, 3:05 PM UTC
HN karma
1,662
Public activity
237 items

About mxwsn

No profile information was provided.

Recent public activity

  1. comment
    Comment #48884181

    This is the same reasoning behind why Yann Lecun thought test-time scaling would not work for LLMs: compounding error. Instead, the more tokens LLMs use, the better their performan…

  2. story
  3. comment
    Comment #48248956

    No, there are more training tokens than parameters in LLMs. They are in the classical first descent setting.

  4. comment
    Comment #48072268

    > Here’s a thought experiment: suppose that a mathematician solved a major problem by having a long exchange with an LLM in which the mathematician played a useful guiding role but…

  5. comment
    Comment #48056545

    Great summary. The fact that the auto encoding task is not grounded in thoughts, and their initial training on guessed internal thoughts, raise serious concerns on faithfulness. Fe…

  6. comment
  7. comment
    Comment #48041064

    Diffusion and flow matching models generate samples by iterative denoising. Iterative denoising means passing input to the neural network, running a forward pass, and taking the ou…

  8. comment
    Comment #47912574

    How do you know that width scaling has been the driving force of improvement?

  9. comment
    Comment #46640794

    The Jacobian is first derivatives, but for a function mapping N to M dimensions. It's the first derivative of every output wrt every input, so it will be an N x M matrix. The gradi…

  10. comment
    Comment #45511981

    Wow! The title suggests introductory material, but in my opinion this has strong potential to win test of time awards for research.

  11. comment
    Comment #45434262

    That's really interesting. What if they RAG search related videos from the prompt, and condition on that to generate? That might explain fidelity like this

  12. comment
    Comment #45342496

    Why is not the diffusion training objective? The technique is known as self-conditioning right? Is it an issue with conditional Tweedie's?

  13. comment
    Comment #44919716

    AI with ability but without responsibility is not enough for dramatic socioeconomic change, I think. For now, the critical unique power of human workers is that you can hold them r…

  14. comment
    Comment #44578257

    Has anyone come across any really cool artifacts? I'd be curious to see

  15. comment
    Comment #44486617

    Stablecoins transferred $27 trillion in 2024 - more than Visa and Mastercard combined. This is right in the article. Stablecoins operate using decentralized ledgers on e.g. Ethereu…

  16. comment
    Comment #44063937

    Gemini has beat it already, but using a different and notably more helpful harness. The creator has said they think harness design is the most important factor right now, and that …

  17. comment
    Comment #43606140

    Huh, I imagined this was because of relaxing regulation.

  18. comment
    Comment #43391115

    Good read, thanks for sharing

  19. comment
    Comment #43287945

    > But what is the original purpose of AI research? I will speak for myself here, but I know many other AI researchers will say the same: the ultimate goal is to understand how huma…

  20. comment
    Comment #43052445

    This ought to be called the qwerty effect, for how the qwerty keyboard layout can't be usurped at this point. It was at the right place at the right time, even though arguably its …

  21. comment
    Comment #42885367

    What's surprising about this is how sparsely defined the rewards are. Even if the model learns the formatting reward, if it never chances upon a solution, there isn't any feedback/…

  22. comment
    Comment #42862418

    I used sublime from 2013 to 2021. It was great. Since, I've switched to VS Code and haven't looked back.

  23. comment
    Comment #42269336

    Context is a challenge for LLMs, but the challenge feels of a different quality to me, than the challenge of incorporating local context into automated decision-making AI like algo…

  24. comment
    Comment #42132502

    My interest was piqued, but the extrapolation in [1] is uh... not the most convincing. If there were more data points then sure, maybe

  25. comment
    Comment #41541148

    OK - there's always a nonzero chance of hallucination. There's also a non-zero chance that macroscale objects can do quantum tunnelling, but no one is arguing that we "need to live…