Live data from Hacker News

GPT-4 gets a B on my quantum computing final exam

scottaaronson.blog

101–110 of 261 posts

Re: GPT-4 gets a B on my quantum computing final exam

#101
post #100
post #73

Reading stuff like this, one thing I cannot stop wondering is this: If Ai can be trusted to do all the trivial tasks and if non trivial tasks require a scaffold of trivial practice, where are we going to keep finding the people qualified enough to actually do the non trivial stuff?

You won't need people for that once AI systems start improving themselves.

So who will then certify them?

Re: GPT-4 gets a B on my quantum computing final exam

#102
post #21
post #6

When I was growing up in the 2000s, it was required to learn a foreign language in school. I took Spanish but dropped the class after a year, thinking that when I grew up, computers would be able to translate text far better humans could. Would you believe it, transformers were invented ten years later. I wonder if it's worth learning anything anymore.

I truly pity you for thinking that learning a foreign language is a redundant exercise because of machine translation. And besides, though machines may perform well on menus and tax returns, I hardly think them on the cusp of emitting fine translations of great poems or novels.

agreed, not to mention that there's things that just can't be expressed. Speaking as someone fluent in spanish and english there's nuance to some things that don't really translate well.

Re: GPT-4 gets a B on my quantum computing final exam

#103

Earlier quoted context omitted.

I took the graduate version of this course, and thankfully did better on this exam than GPT-4. It was a fairly rigorous course, and getting an A required about 20 hours per week of work.

> ... required about 20 hours per week of work. Of thinking or of remembering?

There were weekly problem sets. The questions on these tended to be a bit more in depth than the midterm/exam questions.

Re: GPT-4 gets a B on my quantum computing final exam

#104

Look, it's impressive yes or yes but is it that surprising that a system like an LLM does well on exams? What is an exam? It's an attempt to test that you have consumed and understood large amounts of information and can apply it to novel (ish) situations, but in a very 'sandboxed' way that is just text-in-text-out. That's literally what these systems are specialists in doing. Is it maybe a bit like saying a driverle…

It’s impressive and surprising. Imagine going back to the year 2021 and telling the good people of HN that in 2023 AI would be as advanced like it is now. Literally no one would have believed you in 2021 if you did.

I think it’s reasonable to critique anything and everything. People have, historically, been over eager to believe any AI hype. Hence the numerous AI winters.

Re: GPT-4 gets a B on my quantum computing final exam

#105
post #78
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

I think watching the development of driverless cars in the last 15 years has taught a lot of people to be skeptical of 95% solutions. Sometimes you really need that 100% or the solution is practically useless.

I mean this is clearly not the case with LLM. They create value today even though they are not AGI yet.

Re: GPT-4 gets a B on my quantum computing final exam

#106
post #83
post #60

Earlier quoted context omitted.

>No negativity towards AI here. It’s amazing and it’ll change the future. But we need to be careful on the way. Yeah, I suspect a lot of fields will have a similar trajectory to how AI has impacted radiology. It might catch the tumor in 99.9999% of cases, better than any human doctor. But missing a malignant tumor 0.0001% of the time is unacceptable, because it spikes the hospital's malpractice costs. So every single…

Even with the 1 in 10000 false negative rate, I bet someone is doing the cost calculation of risk vs how many hours it would take for a doctor to check 10000 scans. Doctors themselves are not perfect so they may even have a higher error rate.

Give the doctor an AI tool which is fast and 99.999% accurate. Since they have automation now, give them a massive workload, so they can’t reasonably check everything. Now the machine does the work and the doctor is just the fall-guy if it messes up.

Re: GPT-4 gets a B on my quantum computing final exam

#107
post #60
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

>No negativity towards AI here. It’s amazing and it’ll change the future. But we need to be careful on the way. Yeah, I suspect a lot of fields will have a similar trajectory to how AI has impacted radiology. It might catch the tumor in 99.9999% of cases, better than any human doctor. But missing a malignant tumor 0.0001% of the time is unacceptable, because it spikes the hospital's malpractice costs. So every single…

This seems an oversimplication of radiology. Things are not black and white, we are talking years of training on specific subjects to be able to “see” an image. I believe AI will help, but it will need supervision, at the same rime the doctors are going to get trained on the difficult border cases. Also deanonimizing data for training is a big deal. This is not happening any time soon.

Re: GPT-4 gets a B on my quantum computing final exam

#108
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

A lot of humans fake it until they make it. A lot of humans are lazy. A lot of humans are given responsibility of things that they are unqualified for. A lot of humans make mistakes.

The military, for all its funding and all its training and all its planning, has lost multiple nuclear weapons, on American soil.

We are imperfect machines who aspire to build more perfect versions of ourselves, through children and now through AI. By many measures, we’ve succeeded. The progress will almost certainly continue. The question is, when will it be good enough for you to embrace it despite its imperfections?

Re: GPT-4 gets a B on my quantum computing final exam

#109

An LLM being able to pass your graduate-level exam almost certainly means that your exam is bad. Lots of lazy professors do things like this (true/false and multiple choice answers, very soft questions, etc.), and the presence of GPTs should help them understand that this is not sufficient for evaluating someone's knowledge of a highly technical topic.

Scott Aaronson is one of the best professors and researchers in this field in the world.

That doesn't mean anything about the quality of his courses or exams. Lots of world-class professors phone it in on their teaching responsibilities because they don't care. I had several of them in school - usually more prestigious professors had worse courses.

I can guarantee you that ChatGPT cannot do anything that a quantum computing class prepares you to do, aside from passing this final. That makes it a bad test.

ChatGPT, in a sense, is the police on that front.

Re: GPT-4 gets a B on my quantum computing final exam

#110
post #78
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

I think watching the development of driverless cars in the last 15 years has taught a lot of people to be skeptical of 95% solutions. Sometimes you really need that 100% or the solution is practically useless.

I guess the million dollar question is what are the problems where a 95% solution works.
Post reply on HN