Viewing profile — anorwell
anorwell
HN member- Joined
- Sun, Oct 03, 2010, 6:01 PM UTC
- HN karma
- 677
- Public activity
- 99 items
- HN profile
- View on Hacker News ↗
About anorwell
Recent public activity
-
comment
Comment #49134053
I had the opposite reaction. Clear, detailed, well-organized. Pretty close to the ideal writeup.
-
comment
Comment #48958859
It was rejected for being wrong (or most charitably, incomplete).
-
comment
Comment #47557374
> I don't understand why the models being a year or two old now is worth noting as though it's a clear weakness? I do think it's a clear weakness. Capabilities are extremely differ…
-
comment
Comment #47556522
A pastime I have with papers like this is to look for the part in the paper where they say which models they tested. Very often, you find either A) it's a model from one or more ye…
-
comment
Comment #46964881
HN title editorialization completely inaccurate and misleading here.
-
comment
Comment #46460572
What do you think about the METR 50% task length results? About benchmark progress generally?
-
comment
Comment #46264706
https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com... From my perspective, it's not the worst analogy. In both cases, some people were forecasting an exponential trend in…
-
comment
Comment #46185636
The article does not say at any point which model was used. This is the most basic important information when talking about the capabilities of a model, and probably belongs in the…
-
comment
Comment #46139692
But only in the the tip (nightly) build. I'm somewhat tempted to switch to them for this.
-
comment
Comment #45180980
Thanks, makes sense. I found the benchmark src to see it's not fsyncing, so only some of the files will be durable by the time the benchmark is done. The benchmark docs might benef…
-
comment
Comment #45176634
Seems like a really interesting project! I don't understand what's going on with latency vs durability here. The benchmarks [1] report ~1ms latency for sequential writes, but that'…
-
comment
Comment #44806089
I think your example reflects well on oss-20b, not poorly. It (may) show that they've been successful in separating reasoning from knowledge. You don't _want_ your small reasoning …
-
comment
Comment #44723120
Some of the comments so far seem to be misunderstanding this submission. As I understand it: 1. Custom scaffolding (system prompt and tools) using Qwen3-32B achieved 13.75% on Term…
-
comment
Comment #44224427
This actually intersects with two of my current interests. We have, in production, rarely been seeing ThreadPoolExecutor hangs (JDK17) during shutdown. After a lot of debugging, I'…
-
comment
Comment #44147980
Nor does a neuron. Argumentum ad populum, I have the impression that most computer scientists, at least, do not find Searle's argument at all convincing. Too many people for whom G…
-
comment
Comment #44135290
Is it any good? Perhaps we can ask Opus to review it to find out.
-
comment
Comment #44068577
I am arguing (or rather, presenting without argument) that the Chinese room may be conscious, hence calling it a fallacy above. Not that it _is_ conscious, to be clear, but that th…
-
comment
Comment #44068357
> LLM just complete your prompt in a way that match their training data. They do not have a plan, they do not have thoughts of their own. It's quite reasonable to think that LLMs m…
-
comment
Comment #43736137
The article posts a table of latency distributions, but the latencies are simulated based on the assumption that latencies are lognormal. I would be interested to read the article …
-
comment
Comment #40302487
Interestingly, there was exactly one example on the page with three Xes, instead of one, for "extra wrong": > User: What is the MD5 hash of the string "gremlin"? > Assistant: `5d41…
-
comment
Comment #40227601
> Putting ~100% weights on 'heads' is a terrible prediction! For a weighted coin, isn't this the optimal strategy in the absence of other information? `p > p^2 + ( 1 − p )^2`.
-
comment
Comment #39840904
Not available via the claude.ai web UI. At some point I want to experiment with a CLI-based workflow for programming type queries for myself, but there's also significant utility f…
-
comment
Comment #39839576
This has also been my experience so far with a small sample size of side-by-side prompting. Opus has more hallucinations about APIs, fewer correct refusals for things that are not …
-
comment
Comment #39621291
The longevity people think that blood glucose levels are an important predictor for rate of (biological) aging. See e.g. https://www.lifespan.io/topic/blood-glucose-is-a-biomarker-…
-
comment
Comment #36912272
As a CS degree holder, I got two TN visas, both as a Computer Systems Analyst. This was ~10 years ago. I'm curious if something has changed to make Computer Systems Analyst positio…