Live data from Hacker News

Viewing profile — mnicky

mnicky

HN member
Joined
Mon, Jul 04, 2011, 10:01 PM UTC
HN karma
176
Public activity
107 items

About mnicky

No profile information was provided.

Recent public activity

  1. comment
    Comment #49152732

    It’s more that they have a different business case than competing for the top spots on public benchmarks. They seem to be oriented more toward customizing models for the concrete n…

  2. comment
  3. comment
    Comment #49080239

    With current gaps in DNA synthesis screening yes. But this will be improved in the future hopefully.

  4. comment
    Comment #49080159

    > The biorisk scenarios that the AI safety folks flog are fever-dreamed fantasies that have only the most tenuous connection to biological reality. As an expert, could you also pro…

  5. comment
    Comment #49080116

    [flagged]

  6. comment
    Comment #49041482

    Many ways but mostly ordering some service / using others. Either by social engineering, persuasion, paying etc.

  7. comment
    Comment #49024957

    The air gap would probably help and after this incident I hope labs will think about using such a measure when appropriate. On the other hand I think that proper solution for these…

  8. comment
    Comment #49024848

    These days you can only try, that's why I wrote that :) But in the near future labs will be more automated I guess. The other option you can try these days is maybe social engineer…

  9. comment
    Comment #49021089

    That would be something like 70% of their yearly global profit AFAIK.

  10. comment
    Comment #49017475

    I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent …

  11. comment
    Comment #48998764

    That sonds like they can't compete with 3.5 or 3.6 so they must increase the model size and are training v4.

  12. comment
    Comment #48944366

    It's really simple I think. More tokens per same text length means more capacity to encode information. More information means model can potentially perform better. They introduced…

  13. comment
    Comment #48861495

    For things the agent forgets to obey often, at least in Claude Code, there are also "output styles" that are more deeply embedded - into a system prompt - and agent is also periodi…

  14. comment
    Comment #48861459

    In Claude Code there are also "output styles" that are more deeply embedded - into a system prompt - and agent is also periodically reminded of them during the session: https://cod…

  15. comment
    Comment #48854300

    May be related to this from METR evaluation: > GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated

  16. comment
    Comment #48854290

    Well it's smaller model (something like 4T against 10T Fable). So it's faster and cheaper and with a lot of RL and maybe some favorable benchmark selection it can compete on these …

  17. comment
    Comment #48852528

    Well it seems like they removed quite a few 3rd party benchmarks they used for GPT-5.5 release where Opus 4.7 was better and added many new benchmarks created by them where convini…

  18. comment
    Comment #48852424

    "while being more performant" ..on some specific set of benchmarks ;)

  19. comment
    Comment #48852386

    Maybe Terra = mini and Luna = nano?

  20. comment
    Comment #48852255

    Then we are left with what? FrontierCode maybe? IIRC that one evaluates not only if tests pass but also code quality - e.g. whether the maintainer would accept the pull request as …

  21. comment
    Comment #48851257

    This is especially interesting because IIRC the AA benchmark is calibrated so that 1 point and greater difference is statistically significant.

  22. comment
    Comment #48851144

    There's also this: > GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated -- https://www.lesswrong.com/posts/JFjNmPTbH8kL6xtp6/gpt-5-6-th...

  23. comment
    Comment #48850592

    SWE Bench Pro is completely different benchmark than SWE Bench (e.g. Verified) suite was. It only copied the name.

  24. comment
    Comment #48836927

    One angle could be their interpretability research? They understand what's going on in LLMs probably much better than anyone else. This must somehow pay off. I think it's not only …

  25. comment
    Comment #48829193

    My theory is that they don't have Fable-class intelligence so they needed different hype vehicle :) This rename helps build excitement a bit more than just releasing ordinary GPT-5…