Viewing profile — wgd
wgd
HN member- Joined
- Sun, Oct 23, 2011, 8:16 AM UTC
- HN karma
- 373
- Public activity
- 72 items
- HN profile
- View on Hacker News ↗
About wgd
Recent public activity
-
comment
Comment #49219289
If memory serves the DeepInfra offering is marked as fp4 because that's the native precision of the experts (which are of course the majority of the weights in a MoE model) so they…
-
comment
Comment #49152219
There is https://noslopgrenade.com/
-
comment
Comment #48948621
It's actually pretty straightforward to recover file-states from conversation history. I accidentally deleted the wrong repo on my machine once and recreated all the lost work from…
-
comment
Comment #48928644
You don't hear about them much because their models aren't really competitive. I really wanted to try Trinity Large as a daily-driver in the MiniMax M2 sort of niche but I couldn't…
-
comment
Comment #48900401
> DeepSeek and GLM are left out of the tables entirely: we only have rough characters-divided-by-four estimates for them, not real tokenizer counts, and this post is about measured…
-
comment
Comment #48898914
Yeah, I originally expected this to be about a cursive variant which could be plotted as a single-valued function or something.
-
comment
Comment #48887312
A more accurate title might be "Average University Students Can't Identify Czech AI Poetry". The random-chance performance seen here is reminiscent of the 50-50 nonexpert performan…
-
comment
Comment #48877154
At first I thought this excerpt was meant to warn people off without directly alleging AI authorship, but I guess that's less likely since I see you're also the submitter
-
comment
Comment #48848970
The current Deepseek V4 Pro is still just their initial preview AFAIK, with the "real" model release rumored to come later this month. GLM-5.2 might be outperforming simply because…
-
comment
Comment #48848837
Pangram does work, in the specific sense that when it says something was AI authored it is vanishingly unlikely that it was written by a human (who was not deliberately trying to w…
-
comment
Comment #48805515
> The reason that people don't understand why Anthropic wont let the subscription be used with other harnesses Even more specifically, the very fact that people would prefer, if th…
-
comment
Comment #48805337
I've read a lot of his other writings so that context might be informing my reading here but it sounds like he's pretty straightforwardly discussing the potential of aluminum foil …
-
comment
Comment #48800325
Yes, those ones would be at least a somewhat-plausible simulation of a real scenario people care about: a once-clean codebase that was allowed to become messy by a succession of in…
-
comment
Comment #48799868
"agent pipelines that [...] clean a messy [repository]" This feels like a terrible approach, sufficient to condemn the entire study. Apparently half of the "minimal pairs" in this …
-
comment
Comment #48775031
Some dishwashers add a simple timer-based heuristic so if you open it for just a few seconds while you lazily grab something the "clean" indicator stays lit.
-
comment
Comment #48629144
I've always been amazed at how terrible most frontier LLMs are at compaction given how embarrassingly easy it is to come up with half a dozen different RL training evals which woul…
-
comment
Comment #48563246
The problem is that the moment you introduce shared remote hardware there's a slippery slope leading right back down to "just pay an inference host for model tokens". If you're tra…
-
comment
Comment #48523859
[dead]
-
comment
Comment #48522680
I've got a GLM subscription (mostly because I like supporting open model makers, pretty sure my monthly usage is so low that pay-per-token would be more cost effective), so I gener…
-
comment
Comment #48522209
The GLM-5 series is 744B-A40B. This is not a local model for any reasonable definition of local, but it's an open model which means (once they upload the weights in a week or so) t…
-
comment
Comment #48509794
Often in MoE models the experts are quantized while the shared portions, being a much smaller part of the network with greater impact, are kept at higher or full precision. Not fam…
-
comment
Comment #48479413
Yeah, the evidence feature is so terrible that it actively harms the overall reputation of Pangram. The main "is this AI or human?" classification is done with a machine learning m…
-
comment
Comment #48478989
[flagged]
-
comment
Comment #48470552
People say "determinism" but I don't think that's actually the property we care about. For instance you could imagine a compiler that makes heavy use of superoptimization with rand…
-
comment
Comment #48470032
Yeah I agree this is probably outside of the intended scope of the silent sabotage mechanism, but there are plenty of reports of the "loud" safety classifier misfiring on innocuous…