Live data from Hacker News

Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

arxiv.org

31–40 of 62 posts

Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

#31
post #2

This aligns well with my personal experience using gpt-4. The model provides surprisingly good responses on topics which I know are readily available online while being potentially troublesome to find the exact information I want. I have even found it useful when I know there is a tool for what I want but can’t recall the jargon used to find it via Google. Simply describing the rough idea is enough to get the model t…

Similarly, when I think of ChatGPT as a really cool and advanced search engine frontend, its behavior - including its limitations and its failures - make the most sense to me.

Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

#32
post #31
post #2

This aligns well with my personal experience using gpt-4. The model provides surprisingly good responses on topics which I know are readily available online while being potentially troublesome to find the exact information I want. I have even found it useful when I know there is a tool for what I want but can’t recall the jargon used to find it via Google. Simply describing the rough idea is enough to get the model t…

Similarly, when I think of ChatGPT as a really cool and advanced search engine frontend, its behavior - including its limitations and its failures - make the most sense to me.

> a really cool and advanced search engine frontend

This is the saddest version of ChatGPT I can imagine. I found that as search engines emulated natural language, their results got steadily worse.

I just want the Google results and interface from a long time ago.

Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

#33

Earlier quoted context omitted.

What you’re saying gets to the core of why I would call it AGK and not AGI. Training a transformer on known answers to problems and then observing that it can successfully answer questions related to those problems is cheating. I think the way that Ilya suggests that the “test for consciousness is to train a model with an absolute absence of any training example remotely referring to the notion of a self or of feelin…

Meh Intelligence is Intelligence. It's not cheating for people so asserting that it's cheating for machines just seems like goal post shifting more than anything. Like this idea to pass the machines through frankly ridiculous hoops that humans wouldn't even pass is just..ehh. you seen how children with no language development in childhood turn out ? It just misses the point entirely. It's like the user down the threa…

> It's not cheating for people so asserting that it's cheating for machines just seems like goal post shifting more than anything.

I genuinely appreciate this argument, and was considering it myself. In which case, I’d almost argue that we “have” already achieved AGI, and maybe it’s just not that thrilling.

Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

#34
post #2

This aligns well with my personal experience using gpt-4. The model provides surprisingly good responses on topics which I know are readily available online while being potentially troublesome to find the exact information I want. I have even found it useful when I know there is a tool for what I want but can’t recall the jargon used to find it via Google. Simply describing the rough idea is enough to get the model t…

For a significant number of software developers, GPT and Github's Copilot have replaced StackOverflow, and even Googling more generally. It is more than an autocomplete, it is the best resource for software development by far, IMO. It's a tutor that's an expert in virtually every topic.

No, even expert tutors know how to say “I don’t know” in the face of uncertainty, instead of remorselessly spitting up nonsense as language models do.

Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

#35
post #2

This aligns well with my personal experience using gpt-4. The model provides surprisingly good responses on topics which I know are readily available online while being potentially troublesome to find the exact information I want. I have even found it useful when I know there is a tool for what I want but can’t recall the jargon used to find it via Google. Simply describing the rough idea is enough to get the model t…

For a significant number of software developers, GPT and Github's Copilot have replaced StackOverflow, and even Googling more generally. It is more than an autocomplete, it is the best resource for software development by far, IMO. It's a tutor that's an expert in virtually every topic.

I have to completely disagree with this.

Where GPT-4 shines for me is when I have a project swimming around in my head that I want to work on for fun. It can get you off of the ground quickly, and for side projects the quality and correctness of the output isn't that important.

For professional software development, GPT-4 is still wrong way too often for me to feel comfortable using it. And it's not all that much faster than going straight to the source anyways.

Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

#36
post #2

This aligns well with my personal experience using gpt-4. The model provides surprisingly good responses on topics which I know are readily available online while being potentially troublesome to find the exact information I want. I have even found it useful when I know there is a tool for what I want but can’t recall the jargon used to find it via Google. Simply describing the rough idea is enough to get the model t…

> I am immediately let down

Why? I'm not sure how could you expect anything else in the first place.

Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

#37
post #2

This aligns well with my personal experience using gpt-4. The model provides surprisingly good responses on topics which I know are readily available online while being potentially troublesome to find the exact information I want. I have even found it useful when I know there is a tool for what I want but can’t recall the jargon used to find it via Google. Simply describing the rough idea is enough to get the model t…

For a significant number of software developers, GPT and Github's Copilot have replaced StackOverflow, and even Googling more generally. It is more than an autocomplete, it is the best resource for software development by far, IMO. It's a tutor that's an expert in virtually every topic.

I don't agree.

I still use stack overflow regularly as an engineer.

Sometimes GPT-4 will have a quicker tailor-fit answer, but sometimes it will flounder as well.

Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

#38
post #31
post #2

This aligns well with my personal experience using gpt-4. The model provides surprisingly good responses on topics which I know are readily available online while being potentially troublesome to find the exact information I want. I have even found it useful when I know there is a tool for what I want but can’t recall the jargon used to find it via Google. Simply describing the rough idea is enough to get the model t…

Similarly, when I think of ChatGPT as a really cool and advanced search engine frontend, its behavior - including its limitations and its failures - make the most sense to me.

It's a language model, not a search engine. It doesn't work well as one unless integrated into an actual search engine, like Bing does. Without such integration, it's much closer to human memory than search engine - it will recall stuff it has seen many times pretty well and completely fail at stuff it just glanced over once, filling any gaps with made up stuff like a kid on an exam hoping to get at least a few points with their wild guesses.

Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

#39

Earlier quoted context omitted.

Meh Intelligence is Intelligence. It's not cheating for people so asserting that it's cheating for machines just seems like goal post shifting more than anything. Like this idea to pass the machines through frankly ridiculous hoops that humans wouldn't even pass is just..ehh. you seen how children with no language development in childhood turn out ? It just misses the point entirely. It's like the user down the threa…

> It's not cheating for people so asserting that it's cheating for machines just seems like goal post shifting more than anything. I genuinely appreciate this argument, and was considering it myself. In which case, I’d almost argue that we “have” already achieved AGI, and maybe it’s just not that thrilling.

If you define agi to be artificial and generally intelligent at the human level then yes we have.

It seems though that definitions of agi have since shifted to "better than human experts in all tasks" in which case no...not yet.

Re: Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

#40
post #32
post #31

Earlier quoted context omitted.

Similarly, when I think of ChatGPT as a really cool and advanced search engine frontend, its behavior - including its limitations and its failures - make the most sense to me.

> a really cool and advanced search engine frontend This is the saddest version of ChatGPT I can imagine. I found that as search engines emulated natural language, their results got steadily worse. I just want the Google results and interface from a long time ago.

> I found that as search engines emulated natural language, their results got steadily worse

I would wager that that has not been the experience for the general population (read: non-technical people) and/or that degradation of results has not been because of emulating natural language but because of other factors (like advertising dollars).

Search engines have become incredibly more accessible for non-techies during the past 3 decades. Sure, even today a techie is usually able to coax higher quality results out of a search engine, but it's still a pretty recent advancement that an average Joe can just announce their question out loud and a device on the shelf will not only figure out what they are asking with a decent degree of accuracy, but it will also go search for something relevant, extract an answer, and then speak it back to the user in a pretty sensible way.

It is in this senses in particular that ChatGPT feels like a natural progression for search engines.

Post reply on HN