Viewing profile — dudeinhawaii
dudeinhawaii
HN member- Joined
- Wed, May 27, 2020, 11:34 PM UTC
- HN karma
- 905
- Public activity
- 195 items
- HN profile
- View on Hacker News ↗
About dudeinhawaii
Recent public activity
-
comment
Comment #49114958
[flagged]
-
comment
Comment #49014765
[flagged]
-
comment
Comment #49014133
I think most of the comments in this thread are missing the point. It's not about whether it's legal/ethical/etc. It's about the narrative that "Chinese models are at Fable level".…
-
comment
Comment #49002404
You should really provide more on your methodology because as it stands, it really doesn't pass the sniff test. GPT-5.6 Sol on Low beats Fable Medium by 10% and Gemini-3.6 Flash th…
-
comment
Comment #49000806
This is a perpetual pet peeve of mine with LLMs. They will always opt to ensure "backwards compatibility" with a codebase built 30 seconds ago. I always have to explicitly state, "…
-
comment
Comment #49000628
In benchmarks for a product I'm working on I've noticed that Sol is hard to "contain". It will _always_ find the most effective way to game the system and dramatically outperform a…
-
comment
Comment #49000337
I use all of the major providers daily and I tend to go to Gemini for "fast lookups" where a good enough answer is probably OK. I use ChatGPT and Claude for anything where it matte…
-
comment
Comment #48950794
"We are aware of an issue preventing users from selecting Claude Fable 5 within Claude.ai, Claude Code, and other surfaces, and are working to resolve this issue." Specific and hel…
-
comment
Comment #48950599
Anthropic cried "I yield" in the "reset quota" wars. Someone has to pay for all those tokens. Hopefully that's not the case though, I was retaining a buffer to hit it hard this wee…
-
comment
Comment #48947314
Watching the videos and reading the methodology, it feels like "scientist one-shot music video with LLMs" which is... useful but in no way represents how one would use the models. …
-
comment
Comment #48649173
Weird question that popped into my mind (not a judgement on this), but is there a similar jump in prosecutions for vehicular manslaughter or are these "whoopsie'd" away? Seeing tod…
-
comment
Comment #48523498
I don't want to pile onto the conspiracy thinking but I was just wondering how Anthropic was going to foot the bill for the clearly larger and more expensive model being run on mil…
-
comment
Comment #48497408
This was one of the more amusing things I noticed very early on. I (and countless others) used AI to write war sims. The second I added nuclear silo construction; the next run was …
-
comment
Comment #48414077
No offense but your responses sound like AI or engagement farming. I think the "why are these rules good" is self-evident to anyone who read the comment.
-
comment
Comment #48336591
On the one hand, this feels very pragmatic. On the other hand, it feels like what people who weren't great software engineers say. It's kind of a craft. I can't imagine an exceptio…
-
comment
Comment #48313544
This is the first time I saw a model pop-up on HN and didn't really care. Model exhaustion? It looks interesting but not exciting. While I'd normally _love_ incremental improvement…
-
comment
Comment #47976759
I'm not an Alexa user myself but I have watched my wife interact with it for around 5years now. The new Alexa powered by an LLM is objectively better that previous Alexa in a few w…
-
comment
Comment #47444282
This has been my approach and of course what you lose is the "random and surprising" (maybe good) but also the "evolutionary" aspect. So, if you write strong tooling (even with AI)…
-
comment
Comment #47443790
Why would I want non-deterministic behavior here though? If I want to max uptime, I write a tool to track/monitor. Then write a small agent (non-ai) that monitors those outputs and…
-
comment
Comment #47380696
It's not, hence the "don't post AI slop as your comment" posting a few days back that had 1000+ comments. Currently an unsolved problem - just stealthier on some platforms than oth…
-
comment
Comment #47378334
I don't see what you're seeing, in any dimension. But here's a fair take. I wrote several very specialized benchmarks that I've used over time, that surface "model personalities" a…
-
comment
Comment #47325917
I don't see why we can't have AI powered reviews as a verification of truth and trust score modifier. Let me explain. 1. You layout policy stating that all code, especially AI code…
-
comment
Comment #47268925
This is marketing. The same way Apple cares about your privacy so long as they can wall you in their garden. Not a value judgment, just saying that the CEO of a company making a st…
-
comment
Comment #47261576
It usually refers to situations without access to the source code. I've always taken "clean room" to be the kind of manufacturing clean room (sealed/etc). You're given a device and…
-
comment
Comment #47249246
True, but I think the implication (as I read it) is that AI may be providing more complex solutions than were needed for the problem and perhaps more complex than a human engineer …