Viewing profile — mnicky
mnicky
HN member- Joined
- Mon, Jul 04, 2011, 10:01 PM UTC
- HN karma
- 176
- Public activity
- 107 items
- HN profile
- View on Hacker News ↗
About mnicky
No profile information was provided.
Recent public activity
-
comment
Comment #49152732
It’s more that they have a different business case than competing for the top spots on public benchmarks. They seem to be oriented more toward customizing models for the concrete n…
- comment
-
comment
Comment #49080239
With current gaps in DNA synthesis screening yes. But this will be improved in the future hopefully.
-
comment
Comment #49080159
> The biorisk scenarios that the AI safety folks flog are fever-dreamed fantasies that have only the most tenuous connection to biological reality. As an expert, could you also pro…
-
comment
Comment #49080116
[flagged]
-
comment
Comment #49041482
Many ways but mostly ordering some service / using others. Either by social engineering, persuasion, paying etc.
-
comment
Comment #49024957
The air gap would probably help and after this incident I hope labs will think about using such a measure when appropriate. On the other hand I think that proper solution for these…
-
comment
Comment #49024848
These days you can only try, that's why I wrote that :) But in the near future labs will be more automated I guess. The other option you can try these days is maybe social engineer…
-
comment
Comment #49021089
That would be something like 70% of their yearly global profit AFAIK.
-
comment
Comment #49017475
I think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent …
-
comment
Comment #48998764
That sonds like they can't compete with 3.5 or 3.6 so they must increase the model size and are training v4.
-
comment
Comment #48944366
It's really simple I think. More tokens per same text length means more capacity to encode information. More information means model can potentially perform better. They introduced…
-
comment
Comment #48861495
For things the agent forgets to obey often, at least in Claude Code, there are also "output styles" that are more deeply embedded - into a system prompt - and agent is also periodi…
-
comment
Comment #48861459
In Claude Code there are also "output styles" that are more deeply embedded - into a system prompt - and agent is also periodically reminded of them during the session: https://cod…
-
comment
Comment #48854300
May be related to this from METR evaluation: > GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated
-
comment
Comment #48854290
Well it's smaller model (something like 4T against 10T Fable). So it's faster and cheaper and with a lot of RL and maybe some favorable benchmark selection it can compete on these …
-
comment
Comment #48852528
Well it seems like they removed quite a few 3rd party benchmarks they used for GPT-5.5 release where Opus 4.7 was better and added many new benchmarks created by them where convini…
-
comment
Comment #48852424
"while being more performant" ..on some specific set of benchmarks ;)
-
comment
Comment #48852386
Maybe Terra = mini and Luna = nano?
-
comment
Comment #48852255
Then we are left with what? FrontierCode maybe? IIRC that one evaluates not only if tests pass but also code quality - e.g. whether the maintainer would accept the pull request as …
-
comment
Comment #48851257
This is especially interesting because IIRC the AA benchmark is calibrated so that 1 point and greater difference is statistically significant.
-
comment
Comment #48851144
There's also this: > GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated -- https://www.lesswrong.com/posts/JFjNmPTbH8kL6xtp6/gpt-5-6-th...
-
comment
Comment #48850592
SWE Bench Pro is completely different benchmark than SWE Bench (e.g. Verified) suite was. It only copied the name.
-
comment
Comment #48836927
One angle could be their interpretability research? They understand what's going on in LLMs probably much better than anyone else. This must somehow pay off. I think it's not only …
-
comment
Comment #48829193
My theory is that they don't have Fable-class intelligence so they needed different hype vehicle :) This rename helps build excitement a bit more than just releasing ordinary GPT-5…