Viewing profile — janaksunil
janaksunil
HN member- Joined
- Thu, Oct 03, 2019, 6:43 AM UTC
- HN karma
- 2
- Public activity
- 20 items
- HN profile
- View on Hacker News ↗
About janaksunil
No profile information was provided.
Recent public activity
-
comment
Comment #49680959
do you have some time to chat? janak@withspecific.com
-
comment
Comment #49680762
would love to chat and learn more about your set up! here's my email - janak@withspecific.com
-
comment
Comment #49680729
we manually vet all codebases and companies
-
comment
Comment #49680725
yes!
-
comment
Comment #49680722
that's fair - for long horizon engineering tasks would speed still matter?
-
comment
Comment #49680715
this is a good question. what would make you reject an otherwise working PR on design grounds?
-
comment
Comment #49680704
we reached out to companies that were willing to license their codebases. every codebase we used had real users, one of them had 200k+ users and is currently top 100 on the app sto…
-
comment
Comment #49680686
i appreciate the feedback, the benchmark is primarily long horizon real world engineering tasks on big private codebases.
-
comment
Comment #49680675
we're going to open source some of our tasks and model trajectories as well
-
comment
Comment #49680669
all the codebases were written pre-2023, so pre when AI got good at coding
-
comment
Comment #49680665
would love to learn why?
-
comment
Comment #49680664
will do! happy to chat more on janak@withspecific.com as well
-
comment
Comment #49680660
we've done our best to use the native provider's harness. all models were run on 'high' reasoning. this is still v1 and tons of room for improvement - really appreciate your feedba…
-
comment
Comment #49680655
makes sense
-
comment
Comment #49680652
the tasks on the benchmar are long horizon swe tasks - where GLM does surprisingly well
-
comment
Comment #49680645
here's where all the models messed up! its under this section 'Missed requirements are the most common failure' on realswe.withspecific.com we also have the setup in the blog. the …
-
comment
Comment #49680633
nope, these were private codebases
-
comment
Comment #49680629
hey this seems really interesting - what prompted you to test multiple agents on your codebase?
-
comment
Comment #49680620
i'm janak, cofounder of Specific Labs (YC F25) and one of the authors of Real-SWE. if i can help answer any questions please feel free to email me at janak@withspecific.com, happy …
-
story
Show HN: MCP server that finds dev tool credits in your workflow
I made an MCP server that plugs into Claude Code and tells you if there are credits or discounts available for tools in your stack. First version was super spammy - it fired every …