I'm of the opinion that only tests that were designed for students that were allowed to access the internet should be used as a benchmark for LLMs, and this wouldn't be one.
GPT-4 gets a B on my quantum computing final exam
121–130 of 261 posts
Re: GPT-4 gets a B on my quantum computing final exam
#122Earlier quoted context omitted.
Have you actually used GPT-4 for translation? Seriously all this talk about only getting explicit meaning across would be easily dispelled in an afternoon if you only bothered to try.
GPT-4 can (1) translate, (2) plagiarise, and (3) feedback ("thinking out loud"). Its ability to feedback (3) allows it to execute algorithms, but only a certain class of algorithms. Without tailored prompting, it's further restricted to (a weak generalisation of) algorithms spelled out in its corpus. This is very cool, but this is a skill I possess too, so it's rarely useful to me. Its ability to plagiarise (2) can m…
You're starting from weird assumptions that don't hold up on the capabilities of the model and then determining its abilities from there. It's extremely silly. Next time, use a product extensively for the specified task before you declare what it is and isn't good for.
Literally, everything you've said is just wrong. Can't generate "abstract" translations unless overfit. Lol okay. I've translated passages of fiction across multiple novels to test.
Re: GPT-4 gets a B on my quantum computing final exam
#123Earlier quoted context omitted.
Bilingual/Multilingual LLMs are human level translators more or less. The only way you can think "not on the cusp" is if you haven't actually used GPT-4 for translation. Use it and you'll be set straight pretty quickly.
I haven't used GPT-4 for translation so I acknowledge I might be wrong. But GPT-3 was such an irredeemably terrible poet that it made me sceptical that this type of software could ever develop aesthetic taste or artistic vision. Moreover – and I understand this is an uncharitable thing to say, but it is my honest observation – time and again I have noticed the inability of AI cheerleaders to judge literature on its a…
I don't understand how you could possibly have collected enough data to claim this. How many times have you seen an 'AI cheerleader' (whatever that is) attempt to judge the literature on its artistic merits?
Re: GPT-4 gets a B on my quantum computing final exam
#124Earlier quoted context omitted.
I never saw a scantron at my college. I thought UTAustin was an elite school.
STEM at large public schools is almost exclusively 100+ person classes with maybe 2 TAs, who are more interested in their research than grading exams. We end up with easy to grade exams. I went to a similarly, if not more, prestigious public school and cs exams were usually around 20 multiple choice and 2 or so extended response.
Re: GPT-4 gets a B on my quantum computing final exam
#125https://apps.cs.utexas.edu/apps/sites/default/files/classes/...
There course lecture notes were linked to in the blog post, but also:
https://www.scottaaronson.com/qclec.pdf
This is a first course in quantum computing, and doesn't require a background in physics. For what it is, it certainly looks like challenging course, though. University of Texas at Austin must have some good students.
Still, a lot (a majority) of the exam questions were "word problems", or problems where you take a short mathematical step, based on you being familiar with the concepts and definitions. In short, the type of problems where GPT's pattern matching and filling does well. For the few problems where a longer calculation (actually solving a math/physics problem) was required, GPT performed poorly.
Re: GPT-4 gets a B on my quantum computing final exam
#126Re: GPT-4 gets a B on my quantum computing final exam
#127Earlier quoted context omitted.
>No negativity towards AI here. It’s amazing and it’ll change the future. But we need to be careful on the way. Yeah, I suspect a lot of fields will have a similar trajectory to how AI has impacted radiology. It might catch the tumor in 99.9999% of cases, better than any human doctor. But missing a malignant tumor 0.0001% of the time is unacceptable, because it spikes the hospital's malpractice costs. So every single…
>It might catch the tumor in 99.9999% of cases, better than any human doctor. But missing a malignant tumor 0.0001% of the time is unacceptable, because it spikes the hospital's malpractice costs. So every single scan still has to be reviewed manually by a doctor first, then by the AI as a fallback. I find it hard to believe human doctors miss malignant tumors in less than 1 out of every 10 million cases.
Re: GPT-4 gets a B on my quantum computing final exam
#128Earlier quoted context omitted.
I think watching the development of driverless cars in the last 15 years has taught a lot of people to be skeptical of 95% solutions. Sometimes you really need that 100% or the solution is practically useless.
I guess the million dollar question is what are the problems where a 95% solution works.
Re: GPT-4 gets a B on my quantum computing final exam
#129Earlier quoted context omitted.
I haven't used GPT-4 for translation so I acknowledge I might be wrong. But GPT-3 was such an irredeemably terrible poet that it made me sceptical that this type of software could ever develop aesthetic taste or artistic vision. Moreover – and I understand this is an uncharitable thing to say, but it is my honest observation – time and again I have noticed the inability of AI cheerleaders to judge literature on its a…
>time and again I have noticed the inability of AI cheerleaders to judge literature on its artistic merits. This doesn't hold universally, but it's common enough that I have resolved to regard such claims with extreme doubt. I don't understand how you could possibly have collected enough data to claim this. How many times have you seen an 'AI cheerleader' (whatever that is) attempt to judge the literature on its arti…
Re: GPT-4 gets a B on my quantum computing final exam
#130Earlier quoted context omitted.
Have you actually used GPT-4 for translation? Seriously all this talk about only getting explicit meaning across would be easily dispelled in an afternoon if you only bothered to try.
Agreed. A lot of these responses read like they haven't actually tried it yet. Which is also interesting, I myself actively put off trying it until I eventually gave in. It seems a lot of us are doing the same, maybe its a case of "how good could it actually be?"
Dude clearly hasn't used GPT for translation before and his next reply is telling me the ways GPT should fail based on his pre-conceived notions of its abilities. Except i have actually extensively tested(publicly too) LLMs for translation (even before GPT-4) and basically everything he says is just plain wrong.
I'll never understand why people behave like this.