Live data from Hacker News

Benchmarking GPT-4 Turbo – A Cautionary Tale

blog.mentat.ai

1–10 of 118 posts

Re: Benchmarking GPT-4 Turbo – A Cautionary Tale

#4
> We designed a test for this theory: we reran the benchmarks without showing the models the instructions to each exercise. Instead, we just told them that they were Exercism exercises, and gave them the exercise names and function stubs.

This summarizes all my skepticism agains the AI field. Pretty clear that they aren't solving the problem, they have them memorized.

Re: Benchmarking GPT-4 Turbo – A Cautionary Tale

#6
post #4

> We designed a test for this theory: we reran the benchmarks without showing the models the instructions to each exercise. Instead, we just told them that they were Exercism exercises, and gave them the exercise names and function stubs. This summarizes all my skepticism agains the AI field. Pretty clear that they aren't solving the problem, they have them memorized.

LLMs are lossy compression

Re: Benchmarking GPT-4 Turbo – A Cautionary Tale

#7
post #4

> We designed a test for this theory: we reran the benchmarks without showing the models the instructions to each exercise. Instead, we just told them that they were Exercism exercises, and gave them the exercise names and function stubs. This summarizes all my skepticism agains the AI field. Pretty clear that they aren't solving the problem, they have them memorized.

"They memorized all the problems" is not what was found here and still a wrong overcorrection.

Re: Benchmarking GPT-4 Turbo – A Cautionary Tale

#9
post #4

> We designed a test for this theory: we reran the benchmarks without showing the models the instructions to each exercise. Instead, we just told them that they were Exercism exercises, and gave them the exercise names and function stubs. This summarizes all my skepticism agains the AI field. Pretty clear that they aren't solving the problem, they have them memorized.

"They memorized all the problems" is not what was found here and still a wrong overcorrection.

“Gpt-4 has more problems memorized than gpt-4 turbo” was exactly what was found here.

That doesn’t mean it’s only able to solve problems in its training set (tho it’s much better at that obviously.)

Re: Benchmarking GPT-4 Turbo – A Cautionary Tale

#10
post #3

This reflects my anecdotal findings in a very narrow domain: GPT-4 consistently resolves complex SQL from natural language, whereas GPT-4-Turbo is hit-and-miss, and similar to 3.5-Turbo in performance.

But that’s not what the article is saying at all?
Post reply on HN