Viewing profile — sgk284
sgk284
HN member- Joined
- Fri, Mar 14, 2008, 12:03 AM UTC
- HN karma
- 4,387
- Public activity
- 569 items
- HN profile
- View on Hacker News ↗
About sgk284
Alum of Brex, Google, Microsoft, Salesforce, Twitter, Convoy.
Contact: steve@logic.inc
Recent public activity
-
comment
Comment #48939163
I ran them for a very short time!
-
comment
Comment #48938465
Hey HN, one of the side projects my startup maintains is this arena of 52 apps implemented by different models. Inkling is the most exciting model launch we've seen from an America…
- story
-
comment
Comment #48865836
Yea, that's an interesting result as well. The Terra apps don't feel 35% less feature-rich. So it seems quite token efficient.
-
comment
Comment #48865467
Similarly, we updated our model arena (52 apps each built by 26 models) to have GPT 5.6 Sol, Terra, and Luna today: https://arena.logic.inc/ It's really interesting to see the Sol/…
-
comment
Comment #48841094
Awesome - will work on getting those in.
-
comment
Comment #48840933
If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonn…
-
story
Show HN: Homecrew – Share agent skills across your team and keep them in sync
Hi HN, we share a ton of agent skills at my startup¹ and keeping them in sync was a pain, especially the full cross-product of every team member and every agent they use. We built …
- story
- story
-
comment
Comment #46430341
Yep, 100% correct. We're still reviewing and advising on test cases. We also write a PRD beforehand (with the LLM interviewing us!) so the scope and expectations tend to be fairly …
-
comment
Comment #46428363
It doesn't require removing them if you think you'll need them. It just requires writing tests for those edge cases so you have confidence that the code will work correctly if/when…
-
comment
Comment #46427694
FWIW all of the content on our eng blog is good ol' cage-free grass-fed human-written content. (If the analogy, in the first paragraph, of a Roomba dragging poop around the house d…
-
comment
Comment #46426479
I suspect it will still fall on humans (with machine assistance?) to move the field forward and innovate, but in terms of training an LLM on genuinely new concepts, they tend to be…
-
comment
Comment #46426419
I never claim that 100% coverage has anything to do with code breaking. The only claim made is that anything less than 100% does guarantee that some piece of code is not automatica…
-
comment
Comment #46426351
Can you say more? I see a lot of teams struggling with getting AI to work for them. A lot of folks expect it to be a little more magical and "free" than it actually is. So this pos…
- story
- story
- story
- story
- story
-
comment
Comment #46383741
Reranking is definitely the way to go. We personally found common reranker models to be a little too opaque (can't explain to the user why this result was picked) and not quite ste…
- story
- story
- story