Live data from Hacker News

Viewing profile — imenani

imenani

HN member
Joined
Sat, Nov 11, 2017, 3:03 PM UTC
HN karma
87
Public activity
11 items

About imenani

No profile information was provided.

Recent public activity

  1. comment
    Comment #49155253

    For anyone wondering “how slow is this?” IIUC, Kimi K3 on RTX 6000 Ada (48GB) takes 292 s/token https://github.com/lyogavin/airllm/releases/tag/v3.1.0

  2. comment
    Comment #49136884

    The relation to current RLVR methods I think is interesting, they do discuss it a bit but I would be curious to see more about this as well. Quote from the paper: Exploration beyon…

  3. comment
    Comment #48820885

    Nice presentation of the list! I'd recommend watching a few of his talks/podcasts before during reading these to get the overview and how all the bits in these works tie together. …

  4. comment
    Comment #48149300

    https://lwn.net/Articles/1065620/

  5. comment
    Comment #48141172

    https://xcancel.com/tdietterich/status/2055000956144935055

  6. comment
    Comment #48134774

    The author discussed this here four days ago https://news.ycombinator.com/item?id=48077663

  7. comment
    Comment #47706102

    With the benefit of hindsight, perhaps much of this was Claude Mythos? The model was deployed internally since Feb

  8. comment
    Comment #47555922

    Agreed. LLMs have helped me achieve much deeper reading, _when directed to do so_. Asking an LLM to “Teach me Socratically about this paper/code. One question at a time”, usually a…

  9. comment
    Comment #44858678

    Each of these models has a thinking/reasoning variant and a default non-thinking variant. I would expect the reasoning variants (o3 or “GPT5 Thinking”, Gemini DeepThink, Claude wit…

  10. comment
    Comment #44857965

    As far as I can tell they don’t say which LLM they used which is kind of a shame as there is a huge range of capabilities even in newly released LLMs (e.g. reasoning vs not).

  11. comment
    Comment #43775762

    They fix the temperature at T=0.6 for all k for all models, even though their own Figure 10 shows that RL model benefits from higher temperatures. I would buy the overall claim muc…