Viewing profile — stalfie
stalfie
HN member- Joined
- Sat, Oct 25, 2025, 11:33 AM UTC
- HN karma
- 153
- Public activity
- 65 items
- HN profile
- View on Hacker News ↗
About stalfie
No profile information was provided.
Recent public activity
-
comment
Comment #48886111
Oh I wasn't talking about humans. I should probably have pointed that out. There's some scenarios like catastrophic crop failure and so on that might lead to that, but frankly I do…
-
comment
Comment #48882309
Sure, here are some citations: https://arxiv.org/abs/2401.03910?utm_source=chatgpt.com https://www.frontiersin.org/journals/psychology/articles/10.... https://arxiv.org/pdf/2408.04…
-
comment
Comment #48880970
"Climate conspiracy"? Like you, mean, the conspiracy of climate scientists to publish facts to the best of their understanding? I don't know what exact strawman you're arguing agai…
-
comment
Comment #48877358
No one knows how LLMs work. We know how the architecture works, but almost nothing about why. Saying "statistical next token prediction" tells you about as much about LLMs as sayin…
-
comment
Comment #48877230
I was referring more to the fact that no one predicted that next token predicting Transformers would go so far. Not about "AI" in general.
-
comment
Comment #48872350
No one imagined LLMs in their current format, it was simply a result of discovering that scaling compute and tokens produced better and better results with the Transformer architec…
-
comment
Comment #48792895
10 years is a long time. 10 years ago the Transformer architecture didn't exist. I would call it moderately unlikely at best. At the very least, I would say it's likely that develo…
-
comment
Comment #48784888
There's at least one benchmark that attempts to measure this, but it has been running for a year plus so it's quite infrequently updated now. https://fiction.live/stories/Fiction-l…
-
comment
Comment #48611719
Ironically, the article points out that the original authors publisher actually put out two DMCA notices to google last year, apparently with no effect. I guess DMCA takedowns are …
-
comment
Comment #48611533
Mmmm, not sure I agree with this, although this is a topic where we would have to do a lot of groundwork to formulate our positions precisely in order to ensure we're actually disc…
-
comment
Comment #48609418
Once again, Russia turns out to be the reason we can't have nice things. War truly is a waste for everyone involved. Now that Russia is also helping North Korea to launch satellite…
-
comment
Comment #48609353
Well, I'd argue that this depends on the field you're investigating. Sometimes you have a way to identify objective reality and sometimes you don't. In mathematics the majority of …
-
comment
Comment #48609180
I guess so. Just to be clear, I was talking about post-training methods for reasoning models here, not pre-training. I think "model as a judge" should actually do okay as a "sentim…
-
comment
Comment #48608127
One thing I wonder about hallucinations, is that it seems on the surface that it is an easy problem for RLVR to target. Since you're already generating enormous amounts of reasonin…
-
comment
Comment #48588948
The criticism is also similar to those faced by Theranos. Survivorship bias is always a factor when looking backwards.
-
comment
Comment #48588900
Excluding the cost of X-ray/CT/MRI machines, operating them, getting people to them and through them, sometimes injecting contrast, and sometimes dealing with side effects of said …
-
comment
Comment #48582102
It's worse then that unfortunately. Even when invasive tests are positive, and we think we caught a cancer early, we know from population statistics that the reality is that often …
-
comment
Comment #48487120
Well, the Fable guardrails breaks this argument, as when you get booted down to Opus 4.8 it still happily responds (as does most other models after a "I'm not a doctor but..." hedg…
-
comment
Comment #48477953
I got so frustrated with Fable refusing to make any medical diagnoses, which is the most recent iteration of a longer trend that has bothered me for years, that I made a blog to ex…
- story
-
comment
Comment #48476061
Honestly, I have yet to see any evidence of data leak from private sources. I think one of the better example is "simple-bench", which at least used to be a low-key benchmark that …
-
comment
Comment #48473951
Update in case anyone reads this comment ever again. I have found that I trigger the guardrails any time I ask for medical Q&A as a doctor, be it ECGs, case reports, and so on. But…
-
comment
Comment #48467937
Tried to benchmark ECG interpretation capabilities, and I hit the guardrails no matter what I do. Incredibly frustrating that medical performance seems to be a victim of "biologica…
-
comment
Comment #48422686
This article describes how Transformers work, but not really how LLMs work. Explaining the underlying architecture gives you about as much insight into how a modern LLM behaves as …
-
comment
Comment #48370369
Last time I checked thoroughly (roughly two years ago), AI (in the form of small ML models) mostly outperformed radiologists in areas where the gold standard is "one level" above i…