(from Cursor's blog) > Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools. This dataset lets the model learn both from existing software as well as developer-agent interactions, capturing how developers work and how agents interact with their environments. This is what the big money was for. Cursor is the first big player that had rea…
> You use the previous gen model to prepare datasets for the next model iteration.
you can also use a previous gen model to literally generate data for the next gen model. people used to believe that this is a bad idea but it turns out if you create a scaffold which sinks a lot of compute into generating and grading the data the quality turns out great.