30% drop in O1-preview accuracy when Putnam problems are slightly variated
1–10 of 558 posts
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#2Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#3If o1-preview was the start like gpt1, then we should expect generalization to increase quickly.
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#4Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#5Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/
Not hardcoded, I think it's just likely that those problems exist in its training data in some form
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#6Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/
Not hardcoded, I think it's just likely that those problems exist in its training data in some form
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#7Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#8Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/
Not hardcoded, I think it's just likely that those problems exist in its training data in some form
Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#9Re: 30% drop in O1-preview accuracy when Putnam problems are slightly variated
#10Is it just an open secret that the models are currently just being hardcoded for random benchmarks? Seems weird that people would be asking Putnam problems to a chatbot :/
Not hardcoded, I think it's just likely that those problems exist in its training data in some form