Live data from Hacker News

Viewing profile — sgk284

sgk284

HN member
Joined
Fri, Mar 14, 2008, 12:03 AM UTC
HN karma
4,387
Public activity
569 items

About sgk284

CEO @ Logic, Inc (https://logic.inc)

Alum of Brex, Google, Microsoft, Salesforce, Twitter, Convoy.

Contact: steve@logic.inc

Recent public activity

  1. comment
    Comment #48939163

    I ran them for a very short time!

  2. comment
    Comment #48938465

    Hey HN, one of the side projects my startup maintains is this arena of 52 apps implemented by different models. Inkling is the most exciting model launch we've seen from an America…

  3. story
  4. comment
    Comment #48865836

    Yea, that's an interesting result as well. The Terra apps don't feel 35% less feature-rich. So it seems quite token efficient.

  5. comment
    Comment #48865467

    Similarly, we updated our model arena (52 apps each built by 26 models) to have GPT 5.6 Sol, Terra, and Luna today: https://arena.logic.inc/ It's really interesting to see the Sol/…

  6. comment
    Comment #48841094

    Awesome - will work on getting those in.

  7. comment
    Comment #48840933

    If you like this kind of comparison, we have an arena of 52 apps one-shotted across 21 models here: https://arena.logic.inc/ I keep it pretty up to date (tomorrow Grok 4.5 and Sonn…

  8. story
    Show HN: Homecrew – Share agent skills across your team and keep them in sync

    Hi HN, we share a ton of agent skills at my startup¹ and keeping them in sync was a pain, especially the full cross-product of every team member and every agent they use. We built …

  9. story
  10. story
  11. comment
    Comment #46430341

    Yep, 100% correct. We're still reviewing and advising on test cases. We also write a PRD beforehand (with the LLM interviewing us!) so the scope and expectations tend to be fairly …

  12. comment
    Comment #46428363

    It doesn't require removing them if you think you'll need them. It just requires writing tests for those edge cases so you have confidence that the code will work correctly if/when…

  13. comment
    Comment #46427694

    FWIW all of the content on our eng blog is good ol' cage-free grass-fed human-written content. (If the analogy, in the first paragraph, of a Roomba dragging poop around the house d…

  14. comment
    Comment #46426479

    I suspect it will still fall on humans (with machine assistance?) to move the field forward and innovate, but in terms of training an LLM on genuinely new concepts, they tend to be…

  15. comment
    Comment #46426419

    I never claim that 100% coverage has anything to do with code breaking. The only claim made is that anything less than 100% does guarantee that some piece of code is not automatica…

  16. comment
    Comment #46426351

    Can you say more? I see a lot of teams struggling with getting AI to work for them. A lot of folks expect it to be a little more magical and "free" than it actually is. So this pos…

  17. story
  18. story
  19. story
  20. story
  21. story
  22. comment
    Comment #46383741

    Reranking is definitely the way to go. We personally found common reranker models to be a little too opaque (can't explain to the user why this result was picked) and not quite ste…

  23. story
  24. story
  25. story