Viewing profile — bfogelman
bfogelman
HN member- Joined
- Sat, Jul 07, 2018, 8:11 PM UTC
- HN karma
- 11
- Public activity
- 16 items
- HN profile
- View on Hacker News ↗
About bfogelman
No profile information was provided.
Recent public activity
-
comment
Comment #47442712
we’ve been using it internally (on sculptor) and the speed ups are crazy — we can now have our agents run tests all the time and iterate quickly! Excited other people can now give …
-
comment
Comment #47265909
We’re running both vet and codex on all PRs to do code review and have found they compliment each other well. Vet often catches issues that codex does not!
-
comment
Comment #45432154
ooh this is a great idea -- thanks for sharing!
-
comment
Comment #45429572
See this comment for some differences: https://news.ycombinator.com/item?id=45428185
-
comment
Comment #45428673
nope but vibekit looks interesting -- will take a look
-
comment
Comment #45428663
hmm not ideal -- will try and take a look and see whats going wrong
-
comment
Comment #45428648
haha honestly a little bit ya. One key thing we've learned from working on this is that lowering the barrier to working in parallel is key. Making it easy to merge, context switchi…
-
comment
Comment #45428449
in the works! we want it to be possible to always have the best models and agents available
-
comment
Comment #45428433
lffgggg excited to see where you take lingo log :)
-
comment
Comment #45428419
Hopefully in the next couple of days! You can join the discord and we'll post an announcement when its ready https://discord.gg/GvK8MsCVgk
-
comment
Comment #45428056
right now we're using docker -- we're planning to support modal ( https://modal.com/ ) for remote sandboxes and a "local" mode that might use something like worktrees
-
comment
Comment #45427783
Member of the team here, happy to answer questions. Took a lot of ups, downs and work to get here but excited to finally get this out. Even more excited to share other features we'…
-
comment
Comment #37268040
One thing I’d be curious to see is how well this translates to things outside of HumanEval! How does it compare to using ChatGPT for example.
-
comment
Comment #37268012
Glad this work is happening! That said, HumanEval as the current gold standard for benchmarking models is a crime. The dataset itself is tiny (around 150) examples and all the prob…
- story
-
comment
Comment #33795099
If you’re interested there’s a new RL benchmark that was built using Godot (disclaimer I helped make it!) https://github.com/Avalon-Benchmark/avalon