Viewing profile — jumploops
jumploops
HN member- Joined
- Fri, Mar 01, 2019, 9:38 PM UTC
- HN karma
- 2,746
- Public activity
- 554 items
- HN profile
- View on Hacker News ↗
About jumploops
Recent public activity
-
comment
Comment #49202601
[dead]
-
comment
Comment #49191088
I recently spun up a simple app for our annual mango tasting event[0] using Cloudflare Workers and Durable Objects. It worked really well! Excited to see more options outside of Cl…
- story
-
comment
Comment #49185920
Contrary to the title and intro, this appears to be an agentic _workflow_ builder/runner, not an advanced “agent harness” A few things: - they note: “nothing in this post proves it…
-
comment
Comment #48994193
> By the time I was ready to build a keeper, I had accumulated a scar-tissue document that was empirically sufficient to guide an agent through most of the important decisions, at …
-
comment
Comment #48984194
The models are commodities. Valuations, however, are being built on the models themselves as the product.
-
comment
Comment #48916554
Reminder, this is in the context of "dumb human" prompting. The task is to build a MIPS interpreter to run Doom. The "failed" workflow decided that it couldn't prove Doom was booti…
-
comment
Comment #48902453
I’ve thrown my agentic workflow at Terminal Bench 2.1 and it found a bunch of issues (aka failed tests) because the prompts are “bad” and verifiers are overly specific. As an examp…
-
comment
Comment #48838823
All of the benchmarks are pretty terrible when you look under the hood. For context, I've been iterating on a "supervisor" to replace a lot of the rigamarole spent when working wit…
-
comment
Comment #48827770
My preschooler loves the Untitled Goose Game, please vibe-port this to the Switch (:
- story
-
comment
Comment #48784168
Indoor air quality improvements were one of my “pandemic sourdough” activities. After testing a variety of AQI sensors, I ended up acquiring multiple Airthings-branded devices. The…
- story
-
comment
Comment #48741703
We’re in the process of open-sourcing a few sub-projects within a monorepo, and didn’t know this existed! I’m curious what downsides folks have experienced with this tool? Any tips…
-
comment
Comment #48693223
I don't disagree, we've seen performance shift with capacity changes in the past. With that said, I doubt OpenAI would choose to publish a singular coding benchmark for a new model…
-
comment
Comment #48693198
Don't appreciate the slander, but I'll respond anyhow. Contrary to your predisposition, we're actually quite peeved that we might be seeing results from 5.6 instead of 5.5, as it's…
-
comment
Comment #48692964
[dead]
-
comment
Comment #48691353
If you used GPT-5.5 over the last 24 hours or so, you may have already had access to 5.6. I've been running some tests on a harness we're building, and suddenly saw a jump in a few…
-
comment
Comment #48680139
As an American with mostly Western European ancestors (according to a popular DNA testing site), I've always considered Romans as some distant/tangentially related group. It was su…
- story
-
comment
Comment #48510866
I've been quite impressed with DeepSeek v4 Flash running via antirez's ds4[0]. It feels like a GPT-4 class model in terms of "stored knowledge" but is better at long-horizon tool c…
-
comment
Comment #48497444
> the conversation that generates the code is becoming the true source of our software This is close, but not quite spot on. I've found that I'll test more ideas _with code_ using …
-
comment
Comment #48467040
This is neat! I love that your Step 15 shows an accurate version of the 3d helix, rather than the highly-viral "vortex" animation from a few years back[0] It'd be awesome to scale …
-
comment
Comment #48466340
It's interesting that we're seeing these gains when it seems Mythos/Fable is "just" a scaled up version of their existing architecture[0]. When GPT 4.5 launched, the gains compared…
-
comment
Comment #48455927
No connection, just found it posted elsewhere and thought it was interesting!