I think it's interesting that they've benchmarked it against an array of standardized tests. Seems like LLMs would be particularly well suited to this kind of test by virtue of it being simple prompt:response, but I have to say...those results are terrifying. Especially when considering the rate of improvement. bottom 10% to top 10% of LSAT in What are the implications for society when general thinking, reading, and…
GPT-4
111–120 of 1001 posts
Re: GPT-4
#112This technology has been a true blessing to me. I have always wished to have a personal PhD in a particular subject whom I could ask endless questions until I grasped the topic. Thanks to recent advancements, I feel like I have my very own personal PhDs in multiple subjects, whom I can bombard with questions all day long. Although I acknowledge that the technology may occasionally produce inaccurate information, the…
Re: GPT-4
#113Wow, calculus from 1 to 4, and LeetCode easy from 12 to 31; at this rate, GPT-6 will be replacing / augmenting middle/high school teachers in most courses.
Efficiency seeking players will adopt this quickly but self-sustaining bureaucracy has avoided most modernization successfully over the past 30 years - so why not also AI.
Re: GPT-4
#114AGI is a distraction.
The immediate problems are elsewhere: increasing agency and augmented intelligence are all that is needed to cause profound disequilibrium.
There are already clear and in-the-wild applications for surveillance, disinformation, data fabrication, impersonation... every kind of criminal activity.
Something to fear before AGI is domestic, state, or inter-state terrorism in novel domains.
A joke in my circles the last 72 hours? Bank Runs as a Service. Every piece exists today to produce reasonably convincing video and voice impersonations of panicked VC and dump them on now-unmanaged Twitter and TikTok.
If God-forbid it should ever come to cyberwarfare between China and US, control of TikTok is a mighty weapon.
Re: GPT-4
#115Wow, calculus from 1 to 4, and LeetCode easy from 12 to 31; at this rate, GPT-6 will be replacing / augmenting middle/high school teachers in most courses.
I work in math for the first year of the university in Argentina. We have non mandatory take home exercises in each class. If I waste 10 minutes writing them down in the blackboard instead of handing photocopies, I get like the double of answers by students. It's important that they write the answers and I can comment them, because otherwise they get to the midterms and can't write the answers correctly or they are just wrong and didn't notice. So I waste those 10 minutes. Humans are weird and for some task they like another human.
Re: GPT-4
#116I think it's interesting that they've benchmarked it against an array of standardized tests. Seems like LLMs would be particularly well suited to this kind of test by virtue of it being simple prompt:response, but I have to say...those results are terrifying. Especially when considering the rate of improvement. bottom 10% to top 10% of LSAT in What are the implications for society when general thinking, reading, and…
You can see the limitations by comparing e.g. a memorisation-based test (AP History) with one that actually needs abstraction and reasoning (AP Physics).
Re: GPT-4
#117They trumpet the exam results, but isn't it likely that the model has just memorized the exam?
Re: GPT-4
#118I think it's interesting that they've benchmarked it against an array of standardized tests. Seems like LLMs would be particularly well suited to this kind of test by virtue of it being simple prompt:response, but I have to say...those results are terrifying. Especially when considering the rate of improvement. bottom 10% to top 10% of LSAT in What are the implications for society when general thinking, reading, and…
To address your specific comments:
> What are the implications for society when general thinking, reading, and writing becomes like Chess?
This is a profound and important question. I do think that by “general thinking” you mean “general reasoning”.
> What happens when ALL of our decisions can be assigned an accuracy score?
This requires a system where all human’s decisions are optimized against a unified goal (or small set of goals). I don’t think we’ll agree on those goals any time soon.
Re: GPT-4
#119This is huge: "Rather than the classic ChatGPT personality with a fixed verbosity, tone, and style, developers (and soon ChatGPT users) can now prescribe their AI’s style and task by describing those directions in the 'system' message."
Re: GPT-4
#120This technology has been a true blessing to me. I have always wished to have a personal PhD in a particular subject whom I could ask endless questions until I grasped the topic. Thanks to recent advancements, I feel like I have my very own personal PhDs in multiple subjects, whom I can bombard with questions all day long. Although I acknowledge that the technology may occasionally produce inaccurate information, the…
I don't really know Typescript, so I've been using it a lot to supplement my learning, but I find it really hard to accept any of its answers that aren't straight code examples I can test.