Live data from Hacker News

Viewing profile — tedsanders

tedsanders

HN member
Joined
Wed, Dec 12, 2012, 8:26 AM UTC
HN karma
5,961
Public activity
1,083 items

About tedsanders

http://www.tedsanders.com/about

@sandersted

Recent public activity

  1. story
  2. story
  3. story
  4. comment
    Comment #49040223

    No, they're different models. Knowledge cutoff has been updated.

  5. comment
    Comment #48941977

    I work at OpenAI and I can assure you that if we say we don't train on your data, we don't. I acknowledge that if you don't trust OpenAI, then you may not trust me either. But lyin…

  6. story
  7. comment
    Comment #48866047

    Arena can definitely be benchmaxxed a bit, if you try. The distribution of prompts there is very different than usage by regular coders. E.g., lots of requests for one-shot games f…

  8. comment
    Comment #48864312

    Not entirely fixed yet, but should be rarer with 5.6. Don’t have a quantification, unfortunately.

  9. comment
    Comment #48863179

    > Even worse, it's not a fair comparison: they purposefully just used "adaptive" instead of "max" for Fable. We agree models should be compared on a fair basis. Unfortunately, adap…

  10. comment
    Comment #48849387

    As usual, even though GPT-5.6 is releasing today, the rollout in ChatGPT and Codex will be gradual over many hours so that we can make sure service remains stable for everyone (sam…

  11. comment
    Comment #48838030

    Pointing out problems (e.g., hidden tests that assume narrow implementation details) is much easier than fixing them (e.g., creating tests that work for any possible choice of impl…

  12. comment
    Comment #48828569

    I believe what the post meant to communicate is: - alpha testers will start getting access now - everyone will get access Thursday (barring banned countries / individuals) Historic…

  13. comment
    Comment #48828421

    I work at OpenAI and can confirm that's correct: reasoning tokens are discarded after each new user turn (though not after each message or tool call). Our docs show a diagram here:…

  14. comment
    Comment #48689548

    Unfortunately we're not in a position where we can promise an exact date, but we expect it to take weeks (not days or months). It's the best coding model we've ever trained and we'…

  15. comment
    Comment #48689258

    Yeah, we'll share a lot more details and evals when we can release GPT-5.6 widely. We focused on cyber (and bio) here to help explain why it's being held back for now. We would hav…

  16. comment
    Comment #48453976

    Makes sense, thanks. I suppose error bars are tricky if trying to handle problem-to-problem variance, rubric-to-rubric variance, and run-to-run variance all at once.

  17. comment
    Comment #48453175

    The nonprofit (OpenAI Foundation) owns ~26% of the for-profit, plus some extra warrants. The for-profit (OpenAI Group PBC) is what's filing the S-1 Draft. The OpenAI Foundation als…

  18. comment
    Comment #48452920

    Very cool! So glad to see people building and sharing evals that are better than SWE bench. I'm curious - any particular reason you didn't put error bars on the graphs? Seems like …

  19. comment
  20. story
    An OpenAI model has disproved a central conjecture in discrete geometry

    https://x.com/wtgowers/status/2057175727271800912 , https://xcancel.com/wtgowers/status/2057175727271800912

  21. comment
  22. comment
    Comment #48137586

    What do you mean by this? We don’t train on evals, and if we did I’d quit on the spot. (The loose version of this that’s true is that there may exist eval data contamination in pre…

  23. comment
    Comment #48137505

    Thanks - let me clarify that we don’t switch to lightly quantized models by time of day or when under heavy load either. (I used the adjective heavily because that’s what the origi…

  24. comment
    Comment #48131525

    For what it's worth, I work at OpenAI and I can guarantee you that we don't switch to heavily quantized models or otherwise nerf them when we're under high load. It's true that the…

  25. comment
    Comment #48131505

    FYI, Elo isn't an acronym - it's a person's name. No need to capitalize it as ELO.