Live data from Hacker News

Viewing profile — axiom92

axiom92

HN member
Joined
Sun, Feb 08, 2015, 4:41 PM UTC
HN karma
550
Public activity
204 items

About axiom92

https://madaan.github.io

Recent public activity

  1. comment
    Comment #48856494

    What do you think about Grok 4.5 in comparison to Muse Spark 1.1?

  2. comment
    Comment #46100870

    https://www.reddit.com/r/ChatGPT/comments/14sqcg8/anyone_els...

  3. comment
    Comment #46100866

    This was integrated in gpt4 2 years ago: https://www.reddit.com/r/ChatGPT/comments/14sqcg8/anyone_els...

  4. comment
    Comment #46100075

    And even before this work, there was "PAL: Program-aided Language Models" ( https://arxiv.org/abs/2211.10435 , https://reasonwithpal.com/ ). Afaik PaLM (Google's OG big models) tri…

  5. comment
  6. comment
    Comment #45663292

    You can do this at grok.com. There is a "start thread" option below every conversation. You can also read the responses aloud (helpful if you want to do something async).

  7. comment
    Comment #45647380

    Yeah, that's the first formal reference I remember as well (although, BERT is probably the first thing NLP folks will think of after reading about diffusion). I collected a few oth…

  8. comment
    Comment #45096937

    From last neurips https://automix-llm.github.io/automix/

  9. comment
    Comment #44569251

    The demo was done live (as was everything else).

  10. comment
    Comment #39109672

    Looks pretty cool https://www.youtube.com/watch?v=mogSbMD6EcY Although, it seems it's only going to cover the first book (which makes sense, given how difficult the other two would…

  11. comment
    Comment #38922593

    Welcome to one of the most hated parts of the academia.

  12. comment
    Comment #38315849

    The joke is that he doesn't own any OpenAI shares. [1] https://www.cnbc.com/2023/03/24/openai-ceo-sam-altman-didnt-...

  13. comment
    Comment #37935058

    Right, but no separate image encoder + half the size could be very helpful for many applications.

  14. comment
    Comment #37498090

    > tasks that need deterministic outputs and the thing you need to create is already known statically Wow, interesting. Do you have any example for this? I've realized that LLMs are…

  15. comment
    Comment #37141353

    > And basically all servers will have 8xA100 for those wondering: no this is not the norm. My lab at CMU doesn't own any A100s (we have A6000s).

  16. comment
    Comment #36766503

    > The options seemed to be: If I went for it, I’d be penniless, and if I didn’t go for it, I’d be bitter. I’d be bitter going forward. Penniless certainly beats bitter. So I made t…

  17. comment
    Comment #36627840

    Some of our recent/relevant work: https://selfrefine.info/

  18. comment
    Comment #35765593

    https://en.wikipedia.org/wiki/Dark_forest_hypothesis Also a major theme of https://www.amazon.com/Dark-Forest-Remembrance-Earths-Past/d...

  19. comment
    Comment #35507431

    For those curious about self-refining systems: https://selfrefine.info/ (our recent work).

  20. story
  21. comment
    Comment #35246866

    cushman and code-davinci are similar for sure (same architecture). Perhaps that's what they meant.

  22. comment
    Comment #35242564

    Codex (code-davinci-002) is free (limited beta).

  23. comment
    Comment #35242552

    Actually, there is no way to be sure^. If you think about the costs + scale, it's likely to be cushman (code-cushman-001). ^ Unless you are from OpenAI, in which case I have more q…

  24. comment
    Comment #35242543

    codex (code-davinci-002) was free. This is going to be a huge deal for research groups.

  25. comment
    Comment #35065401

    We did some work in exploring why spelling out the rationale before the answer works so well! Talk: https://madaan.github.io/res/presentations/TwoToTango.pdf Paper: https://arxiv.o…