Live data from Hacker News

Comparing Google and ChatGPT

twitter.com

41–50 of 280 posts

Re: Comparing Google and ChatGPT

#41
Yea... when being proactive, in any way that is not adversarial... ChatGTP has shown me that it's capable of providing very specific insights and knowledge when asking about topics Im currently curious about learning. And it works, I learn the type of information I was seeking. When the topics are technical, GPT is very good at crawl, walk, run with things like algorithms. It's great at responding to "well what about...".

Not only do I learn simpler, I gain better communication style myself when figuring out how to communicate with GPT. GPT also has a nice approach for dialog reasoning.

It's filter system may be annoying, however you can easily learn to play GPT's preferred style of knowledge transfer... and it's honestly something we can learn from.

TLDR; IMO ChatGPT expands the concept of learning, and self-tutoring, in an extremely useful way. This is something no search engine of indexed web pages can compete with. Arguably, the utility of index web pages is really degraded for certain types of desired search experiences when compared to ChatGPT... which it seems obv that internet browsing will be eventually incorporated (probably for further reference and narrowed expansion of a topic)

Re: Comparing Google and ChatGPT

#42

It's great, until people realize GPT-3 will generate answers that are demonstrably wrong. (And to make matters worse, can't show/link the source of the incorrect information!)

I just tried Googling "when did the moon explode?" to see if it still gave authoritative answers to bogus questions:

> About an hour after sunset on June 18, 1178, the Moon exploded.

"when did lincoln shoot booth"

> April 14, 1865

Mostly they seem to catch and stop this now, but there was a fun brief period where it was popping up the fact-box for whatever seemed closest to the search terms, so "when did neil armstrong first walk on the earth" would have it confidently assert "21 July 1969".

Re: Comparing Google and ChatGPT

#43
post #8

Earlier quoted context omitted.

Plus, who is going to produce the corpus you feed the magic chat engine?

As everyone starts to adopt AI, are we going to get to a point where the AI is eating itself. I could imagine AI failing similarly to incestuous genetic lines creating mutations.

Yep, as AI starts to get trained on AI-generated data the output may well become unstable, you can't build an infinite motion machine (or an infinite gain machine/infinite SNR amplifier) and the system may degrade to essentially white noise.

Sort of a cyber-kessler syndrome basically. You really don't want AI-generated content in your AI training material, that's actually probably not generating signal for building future models unless it's undergone further refinement that adds value. An artist iterating on AI artwork is adding signal, and a bunch of artist-curated but not iterated AI artworks probably adds a small amount of signal. But un-refined blogspam and trivial "this one looks cool" probably is reducing signal when you consider the overall output, the AI training process is stable and tolerant to a certain degree of AI content but if you fed in a large portion of unrefined second-order/third-order AI content you would probably get a worse overall result.

Watermarking stable diffusion output by default is an extremely smart move in hindsight, although it's trivial to remove, at least people will have to go to the effort of doing so, which will be a small minority of overall users. But it's a bigger problem than that, you can't watermark text really (again, unless it's called out with a "beep boop I am a robot" tag on reddit or similar) and you can already see AI-generated text getting picked up by various places, search engines, etc. This is the "debris is flying around and starting to shatter things" stage of the kessler syndrome.

In the tech world, you already see it with things like those fake review sites that "interpolate" fake results without explicitly calling it out as such... people do them because they're cheap and easy to do at scale and give you an approximation that is reasonable-ish most of the time for hardware configurations that may not be explicitly benched... now imagine that's all content. Wanna search for how to pull a new electrical circuit or fix your washing machine? Could probably be AI generated in the future. Is it right? Maybe...

Untapped sources of true, organic content are going to become unfathomably valuable in the future, and Archive.org is the trillion-dollar gem. Unfortunately, much like tumblr, if anybody actually buys it the lawyers are going to have a fit and make them delete everything and destroy the asset, but, archive has probably the biggest repository of pre-AI organic content on the planet and that is your repo of training material. Probably the only thing remotely comparable is the library of congress or google's scanning project, but those are narrower and focused on specific types of content. You can generally assume almost all content pre-GPT and pre-stable diffusion is organic, but, the amount of generated content is already a significant minority if not the majority of the content. Like the kessler syndrome, you are seeing this proceed quickly, it is hitting mass-adoption within a span of literally a few years and now the stage is primed for the cascade event.

The other implication here is, people probably need to operate in the mindset that there will be an asymptotically bounded amount of provably-organic training content available... it's not so much that in 10 years we will have 100x the content, because a lot of that content can't really be trusted as input material for further training, a lot of it will be second-order content or third-order content generated by bots or AI and that proportion will increase strongly over the next decade. That's not an inherent dealbreaker, but it probably does have implications for what kinds of training regimes you can build next-next-gen models around, the training set is going to be a lot smaller than people imagine, I think.

Re: Comparing Google and ChatGPT

#44
Here's an example on how misleading ChatGPT is:

Prompt:

Can you explain what happens when i enter "ping 16843009" in a linux shell?

Answer:

When you enter the command ping 16843009 in a Linux shell, you are instructing the ping utility to send an Internet Control Message Protocol (ICMP) echo request to the IP address 168.43.9.0. [...]

The correct answer is that it will ping 1.1.1.1

Also ChatGPT missed that fact that 16843009 is bigger than 255 and failed to explain what happens and why.

Re: Comparing Google and ChatGPT

#46
In the replies someone asks a basic physics question.

"In a vacuum do heavier objects fall faster?"

The response from GPT is completely wrong and so confident, it is like an imposter trying to save face.

Re: Comparing Google and ChatGPT

#47

It's great, until people realize GPT-3 will generate answers that are demonstrably wrong. (And to make matters worse, can't show/link the source of the incorrect information!)

Not just that, but the very first example in the thread being used to demonstrate the capabilities of GPT-3 is in fact wrong, or at least way off-topic, and omitted valuable information that Google succinctly includes.

Specifically, GPT-3 tells the asker to use an align* environment to write a differential equation, but this is not usually necessary, and the presence of the * makes it unnumbered, which may not be desirable and anyway isn't made clear. Google included, and GPT-3 omitted, the use of the \partial symbol for a partial differential equation, which while not always necessary, is definitely something I reach for more often than alignment. Furthermore, the statement "This will produce the following output:" should obviously be followed by an image or PDF or something, although that formatting may not be available; it certainly should not be followed by the same source code!

And personally, I usually find that reading a shorter explanation costs less of my mental energy.

Re: Comparing Google and ChatGPT

#48

It's great, until people realize GPT-3 will generate answers that are demonstrably wrong. (And to make matters worse, can't show/link the source of the incorrect information!)

Seems like we could bang in the idea of PageRank in GPT-3 to marginally improve that situation?

Re: Comparing Google and ChatGPT

#50
post #16

Earlier quoted context omitted.

Exactly, i talked to ChatGPT and it gave me a lot of wrong information in an authorative tone. I consider it dangerous as-is.

Turns out humans do this all the time and they actually have real power.

Yes they do, and I do not deny the power of human's ability to confidently spew nonsense.

However, humans do have some known failure cases that help us detect that. For instance, pressing the human on a couple of details will generally show up all but the very best bullshit artists; there is a limit to how fast humans can make crap up. Some of us are decent at the con-game aspects but it isn't too hard to poke through this limit on how fast they can make stuff up.

Computers can confabulate at full speed for gigabytes at a time.

Personally, I consider any GPT or GPT-like technology unsuitable for any application in which truth is important. Full stop. The technology fundamentally, in its foundation, does not have any concept of truth, and there is no obvious way to add one, either after the fact or in its foundation. (Not saying there isn't one, period, but it certainly isn't the sort of thing you can just throw a couple of interns at and get a good start on.)

"The statistically-most likely conclusion of this sentence" isn't even a poor approximation of truth... it's just plain unrelated. That is not what truth is. At least not with any currently even remotely feasible definition of "statistically most likely" converted into math sufficient to be implementable.

And I don't even mean "truth" from a metaphysical point of view; I mean it in a more engineering sense. I wouldn't set one of these up to do my customer support either. AI Dungeon is about the epitome of the technology, in my opinion, and generalized entertainment from playing with a good text mangler. It really isn't good for much else.

Post reply on HN