Live data from Hacker News

Viewing profile — leerob

leerob

HN member
Joined
Sun, Dec 13, 2015, 5:00 AM UTC
HN karma
2,341
Public activity
786 items

About leerob

https://leerob.com

https://meet.hn/city/41.5868654,-93.6249494/Des-Moines

Socials: - x.com/leerob - lee@leerob.com

Recent public activity

  1. comment
    Comment #48766634

    You can read our full technical report here: https://cursor.com/blog/composer-2-technical-report

  2. comment
    Comment #48764364

    (I work at Cursor) We score well on Terminal-Bench and SWE-bench Multilingual. DeepSWE, not so great yet, as it's more for very long-horizon tasks. We're planning to include more p…

  3. comment
    Comment #48764321

    (I work at Cursor) When Composer 2.5 launched, we initially scored very competitively on AA's composite benchmark. I believe 3rd place overall. They have recently updated to use De…

  4. comment
    Comment #48764251

    (I work at Cursor) CursorBench includes many evals from actual engineering tasks from the Cursor team, which include our private codebase. This codebase is held-out from training s…

  5. comment
    Comment #48742021

    Will do. On it.

  6. comment
    Comment #48740610

    (I work at Cursor) Sorry about this, we should have made this more clear. The new privacy mode is needed because we have to store some state to enable running agents in the cloud. …

  7. story
  8. comment
    Comment #47618495

    Glass was a codename while the UI was in early alpha with testers. It redirects to download now because there is no special link anymore. It's just part of Cursor 3 itself.

  9. comment
    Comment #47618388

    I'm an engineer at Cursor, can try to clarify questions here. > I wish they'd keep the old philosophy of letting the developer drive and the agent assist. Even when I'm using AI ag…

  10. comment
    Comment #47459529

    We used a Kimi base, with midtraining and RL on top. Going forward, we'll include the base used in our blog posts, that was a miss. Also, the license is through Fireworks: https://…

  11. comment
    Comment #47442542

    Are there other coding benchmarks we should include next time? We included Teminal-Bench 2.0 and SWE-bench Mulitilingual. We don't plan on reporting SWE-bench Verified, for similar…

  12. comment
    Comment #47364659

    You can disable this if you want, it's under "Inline Diffs" in the Cursor settings.

  13. story
  14. story
  15. comment
    Comment #46951671

    We've found it to be a strong mix of speed and intelligence. It scores higher than Sonnet 4.5 on Terminal-Bench 2, maybe we will post more on this later.

  16. story
  17. comment
    Comment #46914871

    There's a setting to turn this off if you prefer (Cursor Settings > Agent > Attribution).

  18. story
  19. comment
    Comment #46847562

    > in particular the extra layer of diff-review of AI changes (red/green) which is not integrated into git We're making this better very soon! In the coming weeks hopefully.

  20. comment
    Comment #46847555

    (I work at Cursor) We have all these! Plan mode with a GUI + ability to edit plans inline. Todos. A tool for asking the user questions, which will be automatically called or you ca…

  21. story
  22. comment
    Comment #46651346

    > The JS engine used a custom JS VM being developed in vendor/ecma-rs as part of the browser, which is a copy of my personal JS parser project vendored to make it easier to commit …

  23. comment
    Comment #46651337

    Should compile now: https://news.ycombinator.com/item?id=46650998

  24. story
  25. comment
    Comment #46467201

    Hi. I'm an engineer at Cursor. > By prioritizing the vibe coding use case, Cursor made itself unusable for full-time SWEs. This has actually been the opposite direction we're build…