Live data from Hacker News

Viewing profile — srush

srush

HN member
Joined
Fri, May 24, 2013, 5:48 PM UTC
HN karma
449
Public activity
102 items

About srush

http://rush-nlp.com

https://twitter.com/srush_nlp

Recent public activity

  1. comment
    Comment #45755308

    Awesome to hear, I will share with the team.

  2. comment
    Comment #45753097

    Roughly, we had Cursor software engineers record real questions they were asking models, and then had them record the PR that they made that contained the result. We then cleaned t…

  3. comment
    Comment #45753017

    A lot of people use it! It scores very well on our benchmarks, significantly better than Composer-1.

  4. comment
    Comment #45752934

    We use Ray data for our map-style processing jobs. For example one tool have runs over all the rollouts from the RL system and collects qualitative statistics to understand which t…

  5. comment
    Comment #45752628

    Oh good question. Actually speaking at the Ray Summit next week in SF so we will talk more about it. We used Ray throughout the pipeline for running evals, for the RL controller, f…

  6. comment
    Comment #45751775

    We train with a single agent. is that the question?

  7. comment
  8. comment
    Comment #45751747

    Our view is that there is a now a minimal amount of intelligence that is necessary to be productive, and that if you can pair that with speed that is awesome.

  9. comment
    Comment #45750790

    Let us know how you like it.

  10. comment
    Comment #45750784

    Straight up untrue.

  11. comment
    Comment #45750753

    There are lots of good models we like here. But we agree that getting the right point on the smart+fast graph can make agentic coding feel really good. (Cursor researcher)

  12. comment
    Comment #45750729

    Thanks! Yeah, been working here for 9 months now. Fascinated byt agentic coding both as a researcher and user. Totally agree that "smart model" is the table stakes for usefulness t…

  13. comment
    Comment #45750703

    We will do our best. Luckily I don't think there are major telecom companies called Composer-2.

  14. comment
    Comment #45750685

    Unfortunately not, as we used our own internal code for the benchmark. We would also like to see more benchmarks that reflect the day-to-day agentic coding use.

  15. comment
    Comment #45750484

    The blog talks about the training process. Specifically we trained with RL post-training on coding examples.

  16. comment
    Comment #45750463

    Thanks! I do like the labs blog posts as well though, OpenAI and Anthropic have some classics.

  17. comment
    Comment #45750431

    Our primary focus is on RL post-training. We think that is the best way to get the model to be a strong interactive agent.

  18. comment
    Comment #45750267

    We did a lot of internal testing and thought this model was already quite useful for release.

  19. comment
    Comment #45750246

    Cheetah was an earlier (and dumber) version of this model that we used to test production speed. They are both developed in-house. If you liked Cheetah, give this model a try.

  20. comment
    Comment #45750082

    We like the name Composer and were sad to see it go. Excited to bring it back. (Agree Cheetah is a cool name too.)

  21. comment
    Comment #45749761

    There is a footnote that should help with the models. Training is a harder thing to report on, but roughly our finding here is that RL scales.

  22. comment
    Comment #45749734

    We don't wear shoes [1]. [1] https://www.businessinsider.com/no-shoes-policy-in-office-cu...

  23. comment
    Comment #45749609

    Agree that Sonnet 4.5 is an excellent model. Would be curious to hear your experience using Composer though, it's quite good.

  24. comment
    Comment #45749565

    We also are big Tab users here at Cursor. In the blog we talk about the motivation for this project came from thinking about a Tab-like agent.

  25. comment
    Comment #45749526

    Hi everyone, I am an ML researcher at Cursor, and worked on this project. Would love to hear any feedback you may have on the model, and can answer question about the blog post.