Live data from Hacker News

Viewing profile — bfogelman

bfogelman

HN member
Joined
Sat, Jul 07, 2018, 8:11 PM UTC
HN karma
11
Public activity
16 items

About bfogelman

No profile information was provided.

Recent public activity

  1. comment
    Comment #47442712

    we’ve been using it internally (on sculptor) and the speed ups are crazy — we can now have our agents run tests all the time and iterate quickly! Excited other people can now give …

  2. comment
    Comment #47265909

    We’re running both vet and codex on all PRs to do code review and have found they compliment each other well. Vet often catches issues that codex does not!

  3. comment
    Comment #45432154

    ooh this is a great idea -- thanks for sharing!

  4. comment
    Comment #45429572

    See this comment for some differences: https://news.ycombinator.com/item?id=45428185

  5. comment
    Comment #45428673

    nope but vibekit looks interesting -- will take a look

  6. comment
    Comment #45428663

    hmm not ideal -- will try and take a look and see whats going wrong

  7. comment
    Comment #45428648

    haha honestly a little bit ya. One key thing we've learned from working on this is that lowering the barrier to working in parallel is key. Making it easy to merge, context switchi…

  8. comment
    Comment #45428449

    in the works! we want it to be possible to always have the best models and agents available

  9. comment
    Comment #45428433

    lffgggg excited to see where you take lingo log :)

  10. comment
    Comment #45428419

    Hopefully in the next couple of days! You can join the discord and we'll post an announcement when its ready https://discord.gg/GvK8MsCVgk

  11. comment
    Comment #45428056

    right now we're using docker -- we're planning to support modal ( https://modal.com/ ) for remote sandboxes and a "local" mode that might use something like worktrees

  12. comment
    Comment #45427783

    Member of the team here, happy to answer questions. Took a lot of ups, downs and work to get here but excited to finally get this out. Even more excited to share other features we'…

  13. comment
    Comment #37268040

    One thing I’d be curious to see is how well this translates to things outside of HumanEval! How does it compare to using ChatGPT for example.

  14. comment
    Comment #37268012

    Glad this work is happening! That said, HumanEval as the current gold standard for benchmarking models is a crime. The dataset itself is tiny (around 150) examples and all the prob…

  15. story
  16. comment
    Comment #33795099

    If you’re interested there’s a new RL benchmark that was built using Godot (disclaimer I helped make it!) https://github.com/Avalon-Benchmark/avalon