Live data from Hacker News

Viewing profile — anorwell

anorwell

HN member
Joined
Sun, Oct 03, 2010, 6:01 PM UTC
HN karma
677
Public activity
99 items

About anorwell

anorwell.com

Recent public activity

  1. comment
    Comment #49134053

    I had the opposite reaction. Clear, detailed, well-organized. Pretty close to the ideal writeup.

  2. comment
    Comment #48958859

    It was rejected for being wrong (or most charitably, incomplete).

  3. comment
    Comment #47557374

    > I don't understand why the models being a year or two old now is worth noting as though it's a clear weakness? I do think it's a clear weakness. Capabilities are extremely differ…

  4. comment
    Comment #47556522

    A pastime I have with papers like this is to look for the part in the paper where they say which models they tested. Very often, you find either A) it's a model from one or more ye…

  5. comment
    Comment #46964881

    HN title editorialization completely inaccurate and misleading here.

  6. comment
    Comment #46460572

    What do you think about the METR 50% task length results? About benchmark progress generally?

  7. comment
    Comment #46264706

    https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com... From my perspective, it's not the worst analogy. In both cases, some people were forecasting an exponential trend in…

  8. comment
    Comment #46185636

    The article does not say at any point which model was used. This is the most basic important information when talking about the capabilities of a model, and probably belongs in the…

  9. comment
    Comment #46139692

    But only in the the tip (nightly) build. I'm somewhat tempted to switch to them for this.

  10. comment
    Comment #45180980

    Thanks, makes sense. I found the benchmark src to see it's not fsyncing, so only some of the files will be durable by the time the benchmark is done. The benchmark docs might benef…

  11. comment
    Comment #45176634

    Seems like a really interesting project! I don't understand what's going on with latency vs durability here. The benchmarks [1] report ~1ms latency for sequential writes, but that'…

  12. comment
    Comment #44806089

    I think your example reflects well on oss-20b, not poorly. It (may) show that they've been successful in separating reasoning from knowledge. You don't _want_ your small reasoning …

  13. comment
    Comment #44723120

    Some of the comments so far seem to be misunderstanding this submission. As I understand it: 1. Custom scaffolding (system prompt and tools) using Qwen3-32B achieved 13.75% on Term…

  14. comment
    Comment #44224427

    This actually intersects with two of my current interests. We have, in production, rarely been seeing ThreadPoolExecutor hangs (JDK17) during shutdown. After a lot of debugging, I'…

  15. comment
    Comment #44147980

    Nor does a neuron. Argumentum ad populum, I have the impression that most computer scientists, at least, do not find Searle's argument at all convincing. Too many people for whom G…

  16. comment
    Comment #44135290

    Is it any good? Perhaps we can ask Opus to review it to find out.

  17. comment
    Comment #44068577

    I am arguing (or rather, presenting without argument) that the Chinese room may be conscious, hence calling it a fallacy above. Not that it _is_ conscious, to be clear, but that th…

  18. comment
    Comment #44068357

    > LLM just complete your prompt in a way that match their training data. They do not have a plan, they do not have thoughts of their own. It's quite reasonable to think that LLMs m…

  19. comment
    Comment #43736137

    The article posts a table of latency distributions, but the latencies are simulated based on the assumption that latencies are lognormal. I would be interested to read the article …

  20. comment
    Comment #40302487

    Interestingly, there was exactly one example on the page with three Xes, instead of one, for "extra wrong": > User: What is the MD5 hash of the string "gremlin"? > Assistant: `5d41…

  21. comment
    Comment #40227601

    > Putting ~100% weights on 'heads' is a terrible prediction! For a weighted coin, isn't this the optimal strategy in the absence of other information? `p > p^2 + ( 1 − p )^2`.

  22. comment
    Comment #39840904

    Not available via the claude.ai web UI. At some point I want to experiment with a CLI-based workflow for programming type queries for myself, but there's also significant utility f…

  23. comment
    Comment #39839576

    This has also been my experience so far with a small sample size of side-by-side prompting. Opus has more hallucinations about APIs, fewer correct refusals for things that are not …

  24. comment
    Comment #39621291

    The longevity people think that blood glucose levels are an important predictor for rate of (biological) aging. See e.g. https://www.lifespan.io/topic/blood-glucose-is-a-biomarker-…

  25. comment
    Comment #36912272

    As a CS degree holder, I got two TN visas, both as a Computer Systems Analyst. This was ~10 years ago. I'm curious if something has changed to make Computer Systems Analyst positio…