Konwinski Prize
andykonwinski.com
Konwinski Prize
1–10 of 49 posts
Re: Konwinski Prize
#2Re: Konwinski Prize
#3The prizes scale with the model’s score; the total prize pool is between $100,000 and $1,225,000, depending on the top scores.
Re: Konwinski Prize
#4But what I appreciate even more is that we keep pushing the bar for what an AI can/should be able to do. Excited to track this benchmark over time.
Re: Konwinski Prize
#5The Kaggle competition page has more details: https://www.kaggle.com/competitions/konwinski-prize The prizes scale with the model’s score; the total prize pool is between $100,000 and $1,225,000, depending on the top scores.
If your AI can do this, it's worth several orders of magnitude more. Just FYI.
Re: Konwinski Prize
#6Re: Konwinski Prize
#7In a perfect world this wouldn't be necessary, but in the current research environment where benchmarks are the primary currency and are usually taken at face value, more unbiased evals with known methodology but hidden tests are exactly what we need.
Also one reason why, for instance, I trust small but well-curated benchmarks such as Aider (https://aider.chat/docs/leaderboards/) or Wolfram (https://www.wolfram.com/llm-benchmarking-project/index.php.e...) over large, widely targeted, and increasingly saturated or gamed benchmarks such as LMSYS Arena or HumanEval.
Goodhart's law is thriving and it's our duty to fight it.
Re: Konwinski Prize
#8The Kaggle competition page has more details: https://www.kaggle.com/competitions/konwinski-prize The prizes scale with the model’s score; the total prize pool is between $100,000 and $1,225,000, depending on the top scores.
>$1M for the AI that can close 90% of new GitHub issues If your AI can do this, it's worth several orders of magnitude more. Just FYI.
Re: Konwinski Prize
#9Re: Konwinski Prize
#10What would be an example of cheating since it says "no cheating".