Live data from Hacker News

GPT-4 gets a B on my quantum computing final exam

scottaaronson.blog

171–180 of 261 posts

Re: GPT-4 gets a B on my quantum computing final exam

#171
post #7

Earlier quoted context omitted.

I never saw a scantron at my college. I thought UTAustin was an elite school.

STEM at large public schools is almost exclusively 100+ person classes with maybe 2 TAs, who are more interested in their research than grading exams. We end up with easy to grade exams. I went to a similarly, if not more, prestigious public school and cs exams were usually around 20 multiple choice and 2 or so extended response.

Maybe some STEM, maybe lower level courses. I went to a large school and there were maybe 14 people in most of my upper level classes for my degree.

Re: GPT-4 gets a B on my quantum computing final exam

#172

Earlier quoted context omitted.

Bilingual/Multilingual LLMs are human level translators more or less. The only way you can think "not on the cusp" is if you haven't actually used GPT-4 for translation. Use it and you'll be set straight pretty quickly.

I was curious so I asked GPT-4 to translate a bit of french literature, here it is, along with the official translation (I'll let you guess which is which) ------- The tale I'm about to unfold commenced with a mysterious handwriting on an envelope. Within the pen strokes that outlined my name and the address of the Fossil Review, a publication I was associated with and where the letter had been forwarded from, there…

[deleted]

Re: GPT-4 gets a B on my quantum computing final exam

#173

Earlier quoted context omitted.

I was curious so I asked GPT-4 to translate a bit of french literature, here it is, along with the official translation (I'll let you guess which is which) ------- The tale I'm about to unfold commenced with a mysterious handwriting on an envelope. Within the pen strokes that outlined my name and the address of the Fossil Review, a publication I was associated with and where the letter had been forwarded from, there…

assuming the second is GPT. Here is GPT-4's translation when i say, "add a literary flair" to the translation task. The beginning of all that I am about to recount was an unknown script upon an envelope. Within these strokes of ink that traced my name and the address of the Fossil Review, to which I contributed and from where the letter had been forwarded to me, there swirled a blend of ferocity and gentleness. Behin…

Well I actually did try a few prompts to get it the best I could ! But none were as good as the official, although all better than my own best could possibly be.

Re: GPT-4 gets a B on my quantum computing final exam

#174
post #60
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

>No negativity towards AI here. It’s amazing and it’ll change the future. But we need to be careful on the way. Yeah, I suspect a lot of fields will have a similar trajectory to how AI has impacted radiology. It might catch the tumor in 99.9999% of cases, better than any human doctor. But missing a malignant tumor 0.0001% of the time is unacceptable, because it spikes the hospital's malpractice costs. So every single…

ChatGPT has already found an issue with my relative in the ICU that a literal team of doctors and nurses missed. This just happened last week. Unfortunately we checked ChatGPT retroactively after we went through the screw up.

I think people probably overestimate (maybe vastly) how good at differential diagnosis most doctors are.

Re: GPT-4 gets a B on my quantum computing final exam

#175

> To the best of my knowledge—and I double-checked—this exam has never before been posted on the public Internet, and could not have appeared in GPT-4’s training data. Sure, but you can Google the answers to most of the questions. Personally I've accepted that GPT does learn and apply concepts present in its training data, and all of this would be. Learning is part of intelligence but not the whole thing. (I thought…

I used to agree with you. The paper that made me unsure was "Transformers learn in-context by gradient descent" [1]. Basically, the model learns weights that let it run gradient descent at inference time in order to do in context learning. If transformers can learn this, then I think they can learn almost anything given enough compute/parameters/data.

Of course, even if this is true then it's possible that there simply isn't enough high quality data, or that the amount of compute required is beyond our current hardware.

[1] https://arxiv.org/abs/2212.07677

Re: GPT-4 gets a B on my quantum computing final exam

#176

Earlier quoted context omitted.

Bilingual/Multilingual LLMs are human level translators more or less. The only way you can think "not on the cusp" is if you haven't actually used GPT-4 for translation. Use it and you'll be set straight pretty quickly.

I was curious so I asked GPT-4 to translate a bit of french literature, here it is, along with the official translation (I'll let you guess which is which) ------- The tale I'm about to unfold commenced with a mysterious handwriting on an envelope. Within the pen strokes that outlined my name and the address of the Fossil Review, a publication I was associated with and where the letter had been forwarded from, there…

[deleted]

Re: GPT-4 gets a B on my quantum computing final exam

#177
post #143

Earlier quoted context omitted.

I don't think the data supports the fear that AI assisted driving is more or newly dangerous when compared to fully human drivers. Teslas are safer than any other car on the road. Yes they’re newer, but by mile they’re safer. So the fear that “we should be careful” is understandable but ultimately unfounded. We are being careful .

It’s abundantly clear that we are going to expect near-perfect reliability from autonomous vehicles. This isn’t necessarily illogical; they operate in a different context than humans do. We expect humans to make mistakes and we have various ways of dealing with the consequences (eg lawsuits targeted at the responsible individual). The argument from statistics doesn’t appear likely to win the kind of societal approval…

Your point about liability is valid, but it’s far from “abundantly clear” that society needs machines to be near perfect as opposed to just significantly better than humans.

Solve the liability problem and I would 100% take a machine that performs 30% better than a human or helps a human perform 30% better every time because it means fewer humans die on the road.

Re: GPT-4 gets a B on my quantum computing final exam

#178
post #78
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

I think watching the development of driverless cars in the last 15 years has taught a lot of people to be skeptical of 95% solutions. Sometimes you really need that 100% or the solution is practically useless.

Or people need to learn that different problems have different risk profiles. 95% on if I need more eggs is different than driving a 60 MPH vehicle.

Re: GPT-4 gets a B on my quantum computing final exam

#179

Earlier quoted context omitted.

Machine translation isn't super human yet. But yes, it's probably only a few years away.

I don't get why people are so optimistic about machine translation. Computers can get explicit meaning across – that's obvious to anyone who understands linear algebra, information theory, and linguistics. But many aspects of translation (puns, tone, cultural context) aren't just about mapping from one vector space to another. A human, no matter how fluently bilingual, would have to think about the problem, and the c…

[deleted]

Re: GPT-4 gets a B on my quantum computing final exam

#180

Earlier quoted context omitted.

Bilingual/Multilingual LLMs are human level translators more or less. The only way you can think "not on the cusp" is if you haven't actually used GPT-4 for translation. Use it and you'll be set straight pretty quickly.

I was curious so I asked GPT-4 to translate a bit of french literature, here it is, along with the official translation (I'll let you guess which is which) ------- The tale I'm about to unfold commenced with a mysterious handwriting on an envelope. Within the pen strokes that outlined my name and the address of the Fossil Review, a publication I was associated with and where the letter had been forwarded from, there…

The second is profoundly better. The first is unpublishable.
Post reply on HN