Live data from Hacker News

Viewing profile — stalfie

stalfie

HN member
Joined
Sat, Oct 25, 2025, 11:33 AM UTC
HN karma
153
Public activity
65 items

About stalfie

No profile information was provided.

Recent public activity

  1. comment
    Comment #48886111

    Oh I wasn't talking about humans. I should probably have pointed that out. There's some scenarios like catastrophic crop failure and so on that might lead to that, but frankly I do…

  2. comment
    Comment #48882309

    Sure, here are some citations: https://arxiv.org/abs/2401.03910?utm_source=chatgpt.com https://www.frontiersin.org/journals/psychology/articles/10.... https://arxiv.org/pdf/2408.04…

  3. comment
    Comment #48880970

    "Climate conspiracy"? Like you, mean, the conspiracy of climate scientists to publish facts to the best of their understanding? I don't know what exact strawman you're arguing agai…

  4. comment
    Comment #48877358

    No one knows how LLMs work. We know how the architecture works, but almost nothing about why. Saying "statistical next token prediction" tells you about as much about LLMs as sayin…

  5. comment
    Comment #48877230

    I was referring more to the fact that no one predicted that next token predicting Transformers would go so far. Not about "AI" in general.

  6. comment
    Comment #48872350

    No one imagined LLMs in their current format, it was simply a result of discovering that scaling compute and tokens produced better and better results with the Transformer architec…

  7. comment
    Comment #48792895

    10 years is a long time. 10 years ago the Transformer architecture didn't exist. I would call it moderately unlikely at best. At the very least, I would say it's likely that develo…

  8. comment
    Comment #48784888

    There's at least one benchmark that attempts to measure this, but it has been running for a year plus so it's quite infrequently updated now. https://fiction.live/stories/Fiction-l…

  9. comment
    Comment #48611719

    Ironically, the article points out that the original authors publisher actually put out two DMCA notices to google last year, apparently with no effect. I guess DMCA takedowns are …

  10. comment
    Comment #48611533

    Mmmm, not sure I agree with this, although this is a topic where we would have to do a lot of groundwork to formulate our positions precisely in order to ensure we're actually disc…

  11. comment
    Comment #48609418

    Once again, Russia turns out to be the reason we can't have nice things. War truly is a waste for everyone involved. Now that Russia is also helping North Korea to launch satellite…

  12. comment
    Comment #48609353

    Well, I'd argue that this depends on the field you're investigating. Sometimes you have a way to identify objective reality and sometimes you don't. In mathematics the majority of …

  13. comment
    Comment #48609180

    I guess so. Just to be clear, I was talking about post-training methods for reasoning models here, not pre-training. I think "model as a judge" should actually do okay as a "sentim…

  14. comment
    Comment #48608127

    One thing I wonder about hallucinations, is that it seems on the surface that it is an easy problem for RLVR to target. Since you're already generating enormous amounts of reasonin…

  15. comment
    Comment #48588948

    The criticism is also similar to those faced by Theranos. Survivorship bias is always a factor when looking backwards.

  16. comment
    Comment #48588900

    Excluding the cost of X-ray/CT/MRI machines, operating them, getting people to them and through them, sometimes injecting contrast, and sometimes dealing with side effects of said …

  17. comment
    Comment #48582102

    It's worse then that unfortunately. Even when invasive tests are positive, and we think we caught a cancer early, we know from population statistics that the reality is that often …

  18. comment
    Comment #48487120

    Well, the Fable guardrails breaks this argument, as when you get booted down to Opus 4.8 it still happily responds (as does most other models after a "I'm not a doctor but..." hedg…

  19. comment
    Comment #48477953

    I got so frustrated with Fable refusing to make any medical diagnoses, which is the most recent iteration of a longer trend that has bothered me for years, that I made a blog to ex…

  20. story
  21. comment
    Comment #48476061

    Honestly, I have yet to see any evidence of data leak from private sources. I think one of the better example is "simple-bench", which at least used to be a low-key benchmark that …

  22. comment
    Comment #48473951

    Update in case anyone reads this comment ever again. I have found that I trigger the guardrails any time I ask for medical Q&A as a doctor, be it ECGs, case reports, and so on. But…

  23. comment
    Comment #48467937

    Tried to benchmark ECG interpretation capabilities, and I hit the guardrails no matter what I do. Incredibly frustrating that medical performance seems to be a victim of "biologica…

  24. comment
    Comment #48422686

    This article describes how Transformers work, but not really how LLMs work. Explaining the underlying architecture gives you about as much insight into how a modern LLM behaves as …

  25. comment
    Comment #48370369

    Last time I checked thoroughly (roughly two years ago), AI (in the form of small ML models) mostly outperformed radiologists in areas where the gold standard is "one level" above i…