Benchmarking GPT-4 Turbo – A Cautionary Tale
blog.mentat.ai
Benchmarking GPT-4 Turbo – A Cautionary Tale
1–10 of 118 posts
Re: Benchmarking GPT-4 Turbo – A Cautionary Tale
#2Re: Benchmarking GPT-4 Turbo – A Cautionary Tale
#3Re: Benchmarking GPT-4 Turbo – A Cautionary Tale
#4This summarizes all my skepticism agains the AI field. Pretty clear that they aren't solving the problem, they have them memorized.
Re: Benchmarking GPT-4 Turbo – A Cautionary Tale
#5Re: Benchmarking GPT-4 Turbo – A Cautionary Tale
#6> We designed a test for this theory: we reran the benchmarks without showing the models the instructions to each exercise. Instead, we just told them that they were Exercism exercises, and gave them the exercise names and function stubs. This summarizes all my skepticism agains the AI field. Pretty clear that they aren't solving the problem, they have them memorized.
Re: Benchmarking GPT-4 Turbo – A Cautionary Tale
#7> We designed a test for this theory: we reran the benchmarks without showing the models the instructions to each exercise. Instead, we just told them that they were Exercism exercises, and gave them the exercise names and function stubs. This summarizes all my skepticism agains the AI field. Pretty clear that they aren't solving the problem, they have them memorized.
Re: Benchmarking GPT-4 Turbo – A Cautionary Tale
#8Re: Benchmarking GPT-4 Turbo – A Cautionary Tale
#9> We designed a test for this theory: we reran the benchmarks without showing the models the instructions to each exercise. Instead, we just told them that they were Exercism exercises, and gave them the exercise names and function stubs. This summarizes all my skepticism agains the AI field. Pretty clear that they aren't solving the problem, they have them memorized.
"They memorized all the problems" is not what was found here and still a wrong overcorrection.
That doesn’t mean it’s only able to solve problems in its training set (tho it’s much better at that obviously.)
Re: Benchmarking GPT-4 Turbo – A Cautionary Tale
#10This reflects my anecdotal findings in a very narrow domain: GPT-4 consistently resolves complex SQL from natural language, whereas GPT-4-Turbo is hit-and-miss, and similar to 3.5-Turbo in performance.