Live data from Hacker News

Viewing profile — hodgehog11

hodgehog11

HN member
Joined
Wed, Apr 16, 2025, 1:27 AM UTC
HN karma
1,432
Public activity
406 items

About hodgehog11

No profile information was provided.

Recent public activity

  1. comment
    Comment #49175156

    > The game often has its own DRM though which will stop you I think you missed the "know where to look" part. It's called a Steam emulator, for starters. Note that I speak about th…

  2. comment
    Comment #49167817

    Even if you get all of your games via Steam, provided you have them downloaded, you can still run them without Steam if you know where to look. Obviously GOG is far better in this …

  3. comment
    Comment #49091954

    My expertise lies in deep learning theory, and yes, the "intelligence" is coming primarily from scaling up, among other things. There are good reasons for this, but essentially it …

  4. comment
    Comment #49064759

    Agreed, AI is not capable at the moment of coming up with radical ideas to solve the tough problems. Sadly, I would argue many problems in math are likely to be found to be not act…

  5. comment
    Comment #49043752

    Neither Claude nor GPT are acceptable for writing English text. Personally I have found Gemini to be far better, and that is really all I use it for.

  6. comment
    Comment #49043740

    That has always been the major strength of GPT, that's the model you use for checking. It often nearly isn't as good for creation though.

  7. comment
    Comment #49000252

    Agreed. The benchmark closest to my experience is FrontierMath Tier 4. Fable and Sol (90%) are very far ahead of Kimi K3 (not even 40%). Kimi is trained heavily to basic agentic ta…

  8. comment
    Comment #48972549

    They are likely assessing based on "raw intelligence" benchmarks, rather than agentic ones. Fable crushes in those, but that doesn't necessarily translate to microscopic rigor, whi…

  9. comment
    Comment #48972497

    It really does depend on your application. In my domain (math research), it is substantially better. Fable can solve really hard tasks with surprising consistency. It makes mistake…

  10. comment
    Comment #48966338

    Tell that to my colleagues. Despite Sol getting the attention, Fable is really starting to have an impact on mathematicians right now. It has unbelievable insights in a lot of case…

  11. comment
    Comment #48966328

    I think it is pretty safe to say at this point that having large open LLM models available is better for humanity than them remaining proprietary. Echoing Linus Torvalds' recent co…

  12. comment
    Comment #48966313

    The "second only to Fable 5" comment is pretty telling here. I remember early on when a lot of naysayers were saying that Fable was barely an improvement on Opus. Like it or not, A…

  13. comment
    Comment #48966293

    That would go against everything that Dario believes in (note that I refer to the CEO and not the company; the staff at Anthropic are not so ridiculous). He believes in Anthropic b…

  14. comment
    Comment #48966272

    Value models are always going to be there; you can always distill from a larger model. Having a really intelligent model, regardless of the size, is much better for building confid…

  15. comment
    Comment #48966209

    I love convex optimization and there are a few SciML projects I am on where I really need results from there. But in AI research with deep neural networks, it's become a liability,…

  16. comment
    Comment #48966178

    It's not a matter of whether the theory "works"; it's a matter of whether one is asking the right questions. Convex optimization studies how quickly an optimizer can reach the opti…

  17. comment
    Comment #48965134

    No, I have to push back as well, sorry. It takes a very long time to get to the "near-minimizer" stage when training a neural network, and in practice, you never get there (see neu…

  18. comment
    Comment #48958735

    Very confused by this comment. The older (poorer) parts of the ML literature focus on models with convex and (gradient-)Lipschitz objectives, but that's not representative of reali…

  19. comment
    Comment #48885907

    I really envy you. There is a clear divide amongst my colleagues now in terms of who is using Fable and who isn't (this is math work, so it is well and truly better than all the ot…

  20. comment
    Comment #48885868

    For the particular task I'm working on (a mathematical task in validated numerics), even Sol has generally just repeatedly given up. I asked it for the main problems it could not s…

  21. comment
    Comment #48842650

    Overengineering is the name of the game with Fable. Sometimes you don't want that, sometimes you really do, especially as a researcher. It's a very nice tool to have around for tho…

  22. comment
    Comment #48829933

    For some tasks, there is no amount of "steering" that will produce sensible code. The model needs to be sufficiently capable as a baseline; this is the "intent" that people are ref…

  23. comment
    Comment #48829849

    I'm sorry to hear you are unable to use Fable; my partner is in the same boat and it frustrates her immensely to see what I've been able to do with it. As someone who is working wi…

  24. comment
    Comment #48829782

    I would be absolutely stunned if this were really the case in general given how irresponsibly large Fable is, and 5.6 Sol most definitely is not. It depends on what your problems a…

  25. comment
    Comment #48801075

    I don't agree that my argument was a "No true Scotsman", since the argument made above (as far as I read it) was that "all PhDs are a waste of time". My counterargument was that th…