Viewing profile — tedsanders
tedsanders
HN member- Joined
- Wed, Dec 12, 2012, 8:26 AM UTC
- HN karma
- 5,961
- Public activity
- 1,083 items
- HN profile
- View on Hacker News ↗
About tedsanders
@sandersted
Recent public activity
- story
- story
- story
-
comment
Comment #49040223
No, they're different models. Knowledge cutoff has been updated.
-
comment
Comment #48941977
I work at OpenAI and I can assure you that if we say we don't train on your data, we don't. I acknowledge that if you don't trust OpenAI, then you may not trust me either. But lyin…
- story
-
comment
Comment #48866047
Arena can definitely be benchmaxxed a bit, if you try. The distribution of prompts there is very different than usage by regular coders. E.g., lots of requests for one-shot games f…
-
comment
Comment #48864312
Not entirely fixed yet, but should be rarer with 5.6. Don’t have a quantification, unfortunately.
-
comment
Comment #48863179
> Even worse, it's not a fair comparison: they purposefully just used "adaptive" instead of "max" for Fable. We agree models should be compared on a fair basis. Unfortunately, adap…
-
comment
Comment #48849387
As usual, even though GPT-5.6 is releasing today, the rollout in ChatGPT and Codex will be gradual over many hours so that we can make sure service remains stable for everyone (sam…
-
comment
Comment #48838030
Pointing out problems (e.g., hidden tests that assume narrow implementation details) is much easier than fixing them (e.g., creating tests that work for any possible choice of impl…
-
comment
Comment #48828569
I believe what the post meant to communicate is: - alpha testers will start getting access now - everyone will get access Thursday (barring banned countries / individuals) Historic…
-
comment
Comment #48828421
I work at OpenAI and can confirm that's correct: reasoning tokens are discarded after each new user turn (though not after each message or tool call). Our docs show a diagram here:…
-
comment
Comment #48689548
Unfortunately we're not in a position where we can promise an exact date, but we expect it to take weeks (not days or months). It's the best coding model we've ever trained and we'…
-
comment
Comment #48689258
Yeah, we'll share a lot more details and evals when we can release GPT-5.6 widely. We focused on cyber (and bio) here to help explain why it's being held back for now. We would hav…
-
comment
Comment #48453976
Makes sense, thanks. I suppose error bars are tricky if trying to handle problem-to-problem variance, rubric-to-rubric variance, and run-to-run variance all at once.
-
comment
Comment #48453175
The nonprofit (OpenAI Foundation) owns ~26% of the for-profit, plus some extra warrants. The for-profit (OpenAI Group PBC) is what's filing the S-1 Draft. The OpenAI Foundation als…
-
comment
Comment #48452920
Very cool! So glad to see people building and sharing evals that are better than SWE bench. I'm curious - any particular reason you didn't put error bars on the graphs? Seems like …
- comment
-
story
An OpenAI model has disproved a central conjecture in discrete geometry
https://x.com/wtgowers/status/2057175727271800912 , https://xcancel.com/wtgowers/status/2057175727271800912
- comment
-
comment
Comment #48137586
What do you mean by this? We don’t train on evals, and if we did I’d quit on the spot. (The loose version of this that’s true is that there may exist eval data contamination in pre…
-
comment
Comment #48137505
Thanks - let me clarify that we don’t switch to lightly quantized models by time of day or when under heavy load either. (I used the adjective heavily because that’s what the origi…
-
comment
Comment #48131525
For what it's worth, I work at OpenAI and I can guarantee you that we don't switch to heavily quantized models or otherwise nerf them when we're under high load. It's true that the…
-
comment
Comment #48131505
FYI, Elo isn't an acronym - it's a person's name. No need to capitalize it as ELO.