Live data from Hacker News

GPT-4 gets a B on my quantum computing final exam

scottaaronson.blog

131–140 of 261 posts

Re: GPT-4 gets a B on my quantum computing final exam

#131

Earlier quoted context omitted.

GPT-4 can (1) translate, (2) plagiarise, and (3) feedback ("thinking out loud"). Its ability to feedback (3) allows it to execute algorithms, but only a certain class of algorithms. Without tailored prompting, it's further restricted to (a weak generalisation of) algorithms spelled out in its corpus. This is very cool, but this is a skill I possess too, so it's rarely useful to me. Its ability to plagiarise (2) can m…

Ok so you haven't used it then. I don't care about your whack theories on what it can and can't do. I care about results. You're starting from weird assumptions that don't hold up on the capabilities of the model and then determining its abilities from there. It's extremely silly. Next time, use a product extensively for the specified task before you declare what it is and isn't good for. Literally, everything you've…

> Ok so you haven't used it then.

Not only have I used it, I have made several accurate advance predictions about its behaviour and capabilities – some before GPT-4 was even published. I can model these models well enough to fool GPT output detectors into thinking that I am a GPT model. (Give me a writing task that GPT-4 can't be prompted to perform, and I can prove that last fact to you.)

My theories aren't whack. Perhaps I'm not communicating my understanding very well? I'm not saying GPT-4 can't do anything I haven't listed, but that its ability is bounded by what's demonstrated in its corpus (2): the skill is not legitimately due to the model, and you should not expect a GPT-5 to be any better at the tasks. (In fact, it might well be worse: GPT-4 is worse than GPT-3 at some of these things.)

Re: GPT-4 gets a B on my quantum computing final exam

#133
post #52

I'm curious how the predictive aspect of LLMs can generate / "solve" equations. Is it purely "these inputs are most likely followed by this output", and thus needs to have actually seen the problem to get it right, or is it able to infer some of the rules underlying the operations?

Clearly it doesn't need to have seen the exact input before because you can ask GPT-4 to add together two 6 digit numbers and it will get it right.

More impressively, if you generate a random small neural network and then put some samples in the prompt it can do a surprisingly good job of predicting the result of putting other inputs into the network. Presumably doing something that looks a bit like gradient decent, see https://arxiv.org/pdf/2212.07677.pdf

Solely in order to predict the next token, LLMs are learning an incredibly sophisticed model of both computation and the outside world.

Re: GPT-4 gets a B on my quantum computing final exam

#134

Earlier quoted context omitted.

STEM at large public schools is almost exclusively 100+ person classes with maybe 2 TAs, who are more interested in their research than grading exams. We end up with easy to grade exams. I went to a similarly, if not more, prestigious public school and cs exams were usually around 20 multiple choice and 2 or so extended response.

One of my colleagues uses multiple choice exams for his classes (online and in-person) but he gives them a ton of questions. Something like 90 for a 2hr test. His theory was that, sure, you could look up the answers to a few questions, but if you're doing that for every question you'll run out of time.

Your colleague would be wrong. I've passed multiple exams this way (they allowed Internet access because the prof thought the same way).

You can very likely answer around 1/4th of the questions immediately with less than passive participation in the course, so for the remaining 60 questions you'll get two minutes per question, which is completely manageable if you type and read fast enough.

Multiple choice questions are universally terrible and lazy, and should not be more than a small part of the exam mostly meant to provide easy points.

Re: GPT-4 gets a B on my quantum computing final exam

#135
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

How is this any different from using other technology, e.g. a calculator or a power tool? Or a manager and the output of their ICs?

Obviously the scope of what's possible is different but _any_ craftsman using _any_ tool will only be as good as what they verify themselves.

Re: GPT-4 gets a B on my quantum computing final exam

#136

I find it very endearing that Scott opens his final exam with an optional ungraded question about your favorite QM interpretation.

I would very much like to see his analysis on which interpretations correlate with high scores. I would guess the more niche ones do, and Copenhagen and Many-Worlds will do the worst.

Re: GPT-4 gets a B on my quantum computing final exam

#137
post #38
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

I agree with the point you're making here, but it’s also funny that the description of someone passing a test but not being able to do much without a lot of human supervision is… exactly the description of a human college graduate.

It the college graduates who aren’t the way you describe, those who show initiative and responsibility in their work are the best hires. So not much changes.

Re: GPT-4 gets a B on my quantum computing final exam

#138
This just means this exam can be solved by searching internet and compiling the collections in a presentable format.

For human this is not a easy task because first they need to memorized these information, second, they somehow need to not resort to intensive matrix multiplication in order to say something meaningful.

Even the results from human and chatgpt might look similar but they are very different. This is somewhat similar to comparing finding a interpolation polynomial that is close to sin(x) to just save all the values in the considered interval and do a table lookup.

Re: GPT-4 gets a B on my quantum computing final exam

#139
> To the best of my knowledge—and I double-checked—this exam has never before been posted on the public Internet, and could not have appeared in GPT-4’s training data.

Sure, but you can Google the answers to most of the questions. Personally I've accepted that GPT does learn and apply concepts present in its training data, and all of this would be. Learning is part of intelligence but not the whole thing. (I thought this was why, decades ago, most of the research in this general area rebranded itself from the ambitious goal of artificial intelligence to the more humble but accessible goal of machine learning.)

The interesting question now is how much better it can get, in difficulty it can tackle, reliability, and quality of explanations. How long before AI can answer questions whose answers are not widely known, for which templates do not already exist?

Personally, I believe this will not be some trivial matter of simply scaling up more. The waters of intelligence are much deeper than that. It's hard to believe given the rapid one two punch of GPT 3.5 and 4, but we are about to stall.

If I'm wrong, mark this comment and make fun of me in five years. Wrong or right, it's going to be interesting!

Re: GPT-4 gets a B on my quantum computing final exam

#140
post #9
post #5

This is pretty incredible. I guess the thing that we have to keep in mind is that we really can't compete with an AI that understands natural language this well and also has the entirety of the internet as its reference materials. Maybe we need to accept that or rework our definition of what makes humans special

More that most undergrad colleges only teach and test regurgitation because the profs don't care about education.

That isn't true of this course though. I've read the lecture notes, it really is a great introduction to the subject.
Post reply on HN