Live data from Hacker News

Viewing profile — janaksunil

janaksunil

HN member
Joined
Thu, Oct 03, 2019, 6:43 AM UTC
HN karma
2
Public activity
20 items

About janaksunil

No profile information was provided.

Recent public activity

  1. comment
    Comment #49680959

    do you have some time to chat? janak@withspecific.com

  2. comment
    Comment #49680762

    would love to chat and learn more about your set up! here's my email - janak@withspecific.com

  3. comment
    Comment #49680729

    we manually vet all codebases and companies

  4. comment
  5. comment
    Comment #49680722

    that's fair - for long horizon engineering tasks would speed still matter?

  6. comment
    Comment #49680715

    this is a good question. what would make you reject an otherwise working PR on design grounds?

  7. comment
    Comment #49680704

    we reached out to companies that were willing to license their codebases. every codebase we used had real users, one of them had 200k+ users and is currently top 100 on the app sto…

  8. comment
    Comment #49680686

    i appreciate the feedback, the benchmark is primarily long horizon real world engineering tasks on big private codebases.

  9. comment
    Comment #49680675

    we're going to open source some of our tasks and model trajectories as well

  10. comment
    Comment #49680669

    all the codebases were written pre-2023, so pre when AI got good at coding

  11. comment
    Comment #49680665

    would love to learn why?

  12. comment
    Comment #49680664

    will do! happy to chat more on janak@withspecific.com as well

  13. comment
    Comment #49680660

    we've done our best to use the native provider's harness. all models were run on 'high' reasoning. this is still v1 and tons of room for improvement - really appreciate your feedba…

  14. comment
    Comment #49680655

    makes sense

  15. comment
    Comment #49680652

    the tasks on the benchmar are long horizon swe tasks - where GLM does surprisingly well

  16. comment
    Comment #49680645

    here's where all the models messed up! its under this section 'Missed requirements are the most common failure' on realswe.withspecific.com we also have the setup in the blog. the …

  17. comment
    Comment #49680633

    nope, these were private codebases

  18. comment
    Comment #49680629

    hey this seems really interesting - what prompted you to test multiple agents on your codebase?

  19. comment
    Comment #49680620

    i'm janak, cofounder of Specific Labs (YC F25) and one of the authors of Real-SWE. if i can help answer any questions please feel free to email me at janak@withspecific.com, happy …

  20. story
    Show HN: MCP server that finds dev tool credits in your workflow

    I made an MCP server that plugs into Claude Code and tells you if there are credits or discounts available for tools in your stack. First version was super spammy - it fired every …