Viewing profile — srush
srush
HN member- Joined
- Fri, May 24, 2013, 5:48 PM UTC
- HN karma
- 449
- Public activity
- 102 items
- HN profile
- View on Hacker News ↗
About srush
https://twitter.com/srush_nlp
Recent public activity
-
comment
Comment #45755308
Awesome to hear, I will share with the team.
-
comment
Comment #45753097
Roughly, we had Cursor software engineers record real questions they were asking models, and then had them record the PR that they made that contained the result. We then cleaned t…
-
comment
Comment #45753017
A lot of people use it! It scores very well on our benchmarks, significantly better than Composer-1.
-
comment
Comment #45752934
We use Ray data for our map-style processing jobs. For example one tool have runs over all the rollouts from the RL system and collects qualitative statistics to understand which t…
-
comment
Comment #45752628
Oh good question. Actually speaking at the Ray Summit next week in SF so we will talk more about it. We used Ray throughout the pipeline for running evals, for the RL controller, f…
-
comment
Comment #45751775
We train with a single agent. is that the question?
-
comment
Comment #45751762
neat!
-
comment
Comment #45751747
Our view is that there is a now a minimal amount of intelligence that is necessary to be productive, and that if you can pair that with speed that is awesome.
-
comment
Comment #45750790
Let us know how you like it.
-
comment
Comment #45750784
Straight up untrue.
-
comment
Comment #45750753
There are lots of good models we like here. But we agree that getting the right point on the smart+fast graph can make agentic coding feel really good. (Cursor researcher)
-
comment
Comment #45750729
Thanks! Yeah, been working here for 9 months now. Fascinated byt agentic coding both as a researcher and user. Totally agree that "smart model" is the table stakes for usefulness t…
-
comment
Comment #45750703
We will do our best. Luckily I don't think there are major telecom companies called Composer-2.
-
comment
Comment #45750685
Unfortunately not, as we used our own internal code for the benchmark. We would also like to see more benchmarks that reflect the day-to-day agentic coding use.
-
comment
Comment #45750484
The blog talks about the training process. Specifically we trained with RL post-training on coding examples.
-
comment
Comment #45750463
Thanks! I do like the labs blog posts as well though, OpenAI and Anthropic have some classics.
-
comment
Comment #45750431
Our primary focus is on RL post-training. We think that is the best way to get the model to be a strong interactive agent.
-
comment
Comment #45750267
We did a lot of internal testing and thought this model was already quite useful for release.
-
comment
Comment #45750246
Cheetah was an earlier (and dumber) version of this model that we used to test production speed. They are both developed in-house. If you liked Cheetah, give this model a try.
-
comment
Comment #45750082
We like the name Composer and were sad to see it go. Excited to bring it back. (Agree Cheetah is a cool name too.)
-
comment
Comment #45749761
There is a footnote that should help with the models. Training is a harder thing to report on, but roughly our finding here is that RL scales.
-
comment
Comment #45749734
We don't wear shoes [1]. [1] https://www.businessinsider.com/no-shoes-policy-in-office-cu...
-
comment
Comment #45749609
Agree that Sonnet 4.5 is an excellent model. Would be curious to hear your experience using Composer though, it's quite good.
-
comment
Comment #45749565
We also are big Tab users here at Cursor. In the blog we talk about the motivation for this project came from thinking about a Tab-like agent.
-
comment
Comment #45749526
Hi everyone, I am an ML researcher at Cursor, and worked on this project. Would love to hear any feedback you may have on the model, and can answer question about the blog post.