Viewing profile — monoid73
monoid73
HN member- Joined
- Wed, Apr 09, 2025, 1:33 PM UTC
- HN karma
- 8
- Public activity
- 10 items
- HN profile
- View on Hacker News ↗
About monoid73
No profile information was provided.
Recent public activity
-
comment
Comment #44828128
for the hybrid workflows, curious how do you decide which parts need AI reasoning vs can be hardcoded? is it adaptive or manual config?
-
story
Show HN: Open Operator Evals – real-world benchmarks for LLM web agents
We’ve open-sourced a benchmark for LLM-driven web agent setups. It evaluates real-world tasks, like logging in, scraping dashboards, and submitting forms, using structured criteria…
-
comment
Comment #43929546
Another one? People saw that 3B windsurf money.
-
comment
Comment #43811168
think the visa hurdle is the big one. even if you have a strong background, a lot of companies hesitate unless they already have an immigration pipeline set up. another angle could…
-
comment
Comment #43787588
I think the UX of chatgpt works because it's familiar, not because it's good. Lowers friction for new users but doesn't scale well for more complex workflows. if you're building an…
-
comment
Comment #43787567
funny enough, i started noticing em dashes mostly through using GPT. wasn’t really part of my writing before, but now i find them super useful for managing rhythm and flow. definit…
-
comment
Comment #43776331
this is one of the more compelling "LLM meets real-world tool" use cases i've seen. openSCAD makes a great testbed since it's text-based and deterministic, but i wonder what the li…
-
comment
Comment #43766179
exactly. hindsight bias makes it really hard to separate genuine inference from subtle prompt leakage. even framing the question can accidentally steer it toward the right answer. …
-
comment
Comment #43766101
same here. brew’s been great historically but it’s gotten bloated and kinda slow. curious to see if sapphire can keep things lean without sacrificing compatibility.
-
comment
Comment #43757145
yeah, that'd b nice, some kind of self-bootstrapping system where you start with a strong cloud model, then fine-tune a smaller local one over time until it’s good enough to take o…