Live data from Hacker News

GPT-4 gets a B on my quantum computing final exam

scottaaronson.blog

201–210 of 261 posts

Re: GPT-4 gets a B on my quantum computing final exam

#201
post #143

Earlier quoted context omitted.

I don't think the data supports the fear that AI assisted driving is more or newly dangerous when compared to fully human drivers. Teslas are safer than any other car on the road. Yes they’re newer, but by mile they’re safer. So the fear that “we should be careful” is understandable but ultimately unfounded. We are being careful .

It’s abundantly clear that we are going to expect near-perfect reliability from autonomous vehicles. This isn’t necessarily illogical; they operate in a different context than humans do. We expect humans to make mistakes and we have various ways of dealing with the consequences (eg lawsuits targeted at the responsible individual). The argument from statistics doesn’t appear likely to win the kind of societal approval…

> We expect humans to make mistakes and we have various ways of dealing with the consequences (eg lawsuits targeted at the responsible individual).

The same is true of manufacturers and others in the chain of commerce of goods (see, e.g., the general rules on defective product liability), even where they aren’t individual humans. There’s nothing about AI which makes it particularly special in this regard.

Re: GPT-4 gets a B on my quantum computing final exam

#203

Earlier quoted context omitted.

1) It's not possible to fairly compare human intelligence with something that can memorize gigabytes of text and hold it in non-volatile memory. 2) Months ago, in my earliest interactions with ChatGPT, I asked it to solve math problems. It gave me back stuff with LaTeX formatting. Obviously it had, if not these exact problems, similar templates in its training set. Recently it was shown that GPT is completely incapab…

> It gave me back stuff with LaTeX formatting. So? It can write LaTeX just like it can write Java or Python or Rust.

It put LaTeX in there unbidden because it pattern matched to LaTeX source code in its training set. Code containing solved math problems similar to those it was being asked.

Re: GPT-4 gets a B on my quantum computing final exam

#204
post #180

Earlier quoted context omitted.

The second is profoundly better. The first is unpublishable.

Unpublishable ? Huh. I think the second is GPT though neither is unpublishable. i think the first is better for the domain - a novel but the second is more literal. Still though, if you wanted GPT to spit out something like the first or second, you'd only need to ask. "add a literary flair" works (for the first).

Second one is actually the official ! It indeed added some literary flair, but this flair does not have such presence in the original french text, as it is a modern and experimental work of literature.

Re: GPT-4 gets a B on my quantum computing final exam

#205

Earlier quoted context omitted.

GPT4 実際にGPT-4を翻訳に使ったことがありますか?本当に、明示的な意味だけを伝えるという話が、試してみるだけで簡単に解決できるなんて、冗談じゃないですか。

WHAT. It's got the second half wrong. Google: Have you actually used GPT-4 for translation? Really, it's a joke that the story of only conveying explicit meaning can be easily solved by just trying. DeepL: Have you actually used GPT-4 for translation? Really, it's a joke that all this talk about conveying only explicit meaning can be easily solved by just trying it out. Mine: Have you actually used GPT-4 for translat…

Here's a couple more from GPT4 (since it's random every time because of temperature)

GPT-4を翻訳に実際に使ったことがありますか?本気で、伝えたい意味だけを伝えるという話は、ちょっと試してみれば簡単に解決できると思うのですが。

実際にGPT-4を翻訳に使ったことがありますか?本当に、試してみるだけで簡単に払拭できると思うのに、この「明確な意味だけが伝わる」話ばかりで。

Re: GPT-4 gets a B on my quantum computing final exam

#206
post #199
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

It's very easy to stop these AI cold. Make "I don't know" or "Insufficient information" a correct answer--those are not in the training set. These AI have "Male Answer Syndrome" to the nth degree. They will shamelessly give you an answer even if they have to completely make one up.

> Male Answer Syndrome

is that like mansplaining?

is botsplaining the new mansplaining?

Re: GPT-4 gets a B on my quantum computing final exam

#207
post #60
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

>No negativity towards AI here. It’s amazing and it’ll change the future. But we need to be careful on the way. Yeah, I suspect a lot of fields will have a similar trajectory to how AI has impacted radiology. It might catch the tumor in 99.9999% of cases, better than any human doctor. But missing a malignant tumor 0.0001% of the time is unacceptable, because it spikes the hospital's malpractice costs. So every single…

> It might catch the tumor in 99.9999% of cases, better than any human doctor. But missing a malignant tumor 0.0001% of the time is unacceptable

Those probabilities are way off given biology, but anyway ...

The interesting cases of AI in radiology would be being able to catch stuff that a human has no hope of catching.

For example, a woman with lobular (instead of ductal) breast cancer generally doesn't present until mid-to-late Stage 3 (which limits treatment options) because those cancers don't form lumps.

You can stare at mammograms and ultrasounds all day and won't see anything because the "lumps" are unresolvable. You're trying to find a sleet particle in a blizzard. Sure, it's totally obvious on an MRI scan, but you don't want to do those without reason (picking up totally benign growths, gadolinium bioaccumulation, infections from IVs, etc.)

An AI, however, could correlate subtle, but broad changes that humans are really bad at catching. Your last 5 mammograms looked like this but there is just something a little off about this one--go get an MRI this time.

Re: GPT-4 gets a B on my quantum computing final exam

#208
post #38
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

I agree with the point you're making here, but it’s also funny that the description of someone passing a test but not being able to do much without a lot of human supervision is… exactly the description of a human college graduate.

For a college graduate, that is the starting point. Test results are supposed to signal that the person can learn new things. While a fresh graduate needs a lot of supervision, they should quickly become more capable and productive.

For a language model, test results are the end. They are supposed to measure what the model is capable of. If you need better performance, you must train a better model.

Re: GPT-4 gets a B on my quantum computing final exam

#209

Earlier quoted context omitted.

Scott Aaronson is one of the best professors and researchers in this field in the world.

That doesn't mean anything about the quality of his courses or exams. Lots of world-class professors phone it in on their teaching responsibilities because they don't care. I had several of them in school - usually more prestigious professors had worse courses. I can guarantee you that ChatGPT cannot do anything that a quantum computing class prepares you to do, aside from passing this final. That makes it a bad test…

Oh, you can guarantee, huh? Wow, guess that settles it then.

Re: GPT-4 gets a B on my quantum computing final exam

#210
post #165
post #53

Earlier quoted context omitted.

Doubt it … GPTs speak any language they know natively, but if asked for translations they seem unable to deal with sentence structures and logics that exists in the source but not allowed in the target language. When that happens they fail to recognize gibberishness of that.

Can you provide an example? Because from my experience it's the exact opposite - GPT-4 can handle translation, especially when there are complex sentences and context that needs to be kept across sentences, way better than Google Translate currently can.

Maybe this is a dirty secret about Japanese: it's often necessary to basically ghostwrite sentences to go to/from English, because direct word-for-word substitutions won't make sense.

GPTs don't seem to do that, and as far as my exposure to them(<3.5) goes, they don't seem to understand what I'm talking about.

Post reply on HN