Live data from Hacker News

Generative AI could make search harder to trust

wired.com

121–130 of 255 posts

Re: Generative AI could make search harder to trust

#121

I wonder if there will be a human information/knowledge equivalent of low-background steel (pre-WWII/nukes). Data from before a certain point won't be 'contaminated' with LLM stuff, but it'll be everywhere after that. https://en.wikipedia.org/wiki/Low-background_steel

I've said as much as to one extra incentive (besides retrain cost) as to why openai has frozen the training period for post-2022. I think it's trying to generate as much data as possible before itself has contaminated its training set. They'll effectively have a monopoly over it; it's interestingly a rare example of "we can only do this once, then it's forever degraded". You really want a clean & discrete start.

Re: Generative AI could make search harder to trust

#122

I wonder if there will be a human information/knowledge equivalent of low-background steel (pre-WWII/nukes). Data from before a certain point won't be 'contaminated' with LLM stuff, but it'll be everywhere after that. https://en.wikipedia.org/wiki/Low-background_steel

People are vastly everestimating how unique this problem of hallucinations is. It seems to me it relies mostly on discounting just how much we've already had to deal with this same problem in humans over the millenia. The problem of proliferation of bad information might be getting worse, but this isn't native to generative AI. The entire informational ecosystem has to deal with this. GPTs compound the issue, but as…

If there's an upside here, it's that humans will now be forced to refine their BS detectors.

In doing so the species would improve critical thinking skills which can be applied to all information regardless of source. Which, I agree, was often BS to begin with. But in theory would be more difficult to skirt on by without notice if humanity upgraded their critical thinking.

Re: Generative AI could make search harder to trust

#123

Without naming the company, I have seen specific examples of blog posts being written by AI, hallucinating a "fact", and then that "fact" re-surfacing inside of Bard. Its xkcd's Citogenesis automated and at internet scale https://xkcd.com/978/

Someone on Twitter just tried asking Bard [are you familiar with a paper called "A Short History of Searching"] and Bard responded by linking Claude Shannon to a fabrication. https://twitter.com/nunohipolito/status/1710063374145343511

Re: Generative AI could make search harder to trust

#124

I wonder if there will be a human information/knowledge equivalent of low-background steel (pre-WWII/nukes). Data from before a certain point won't be 'contaminated' with LLM stuff, but it'll be everywhere after that. https://en.wikipedia.org/wiki/Low-background_steel

I suspect in the coming years the Wayback Machine at archive.org will become ever more important - always assuming it's not lost as collateral damage in their copyright battles. Indexing that dataset and making it searchable would massively increase its value. My inner conspiracy theorist can't help wonder if the continued reduction in search usefulness isn't part of an ongoing deliberate disempowerment of everyday p…

Generative Wayback Machine

Re: Generative AI could make search harder to trust

#125

I wonder if there will be a human information/knowledge equivalent of low-background steel (pre-WWII/nukes). Data from before a certain point won't be 'contaminated' with LLM stuff, but it'll be everywhere after that. https://en.wikipedia.org/wiki/Low-background_steel

People are vastly everestimating how unique this problem of hallucinations is. It seems to me it relies mostly on discounting just how much we've already had to deal with this same problem in humans over the millenia. The problem of proliferation of bad information might be getting worse, but this isn't native to generative AI. The entire informational ecosystem has to deal with this. GPTs compound the issue, but as…

Not one single solitary soul has ever made the claim that misinformation didn't exist before AI so it's not clear who you're arguing with. People are rightly concerned about the scale of misinformation that AI is unlocking.

Re: Generative AI could make search harder to trust

#126

Earlier quoted context omitted.

People are vastly everestimating how unique this problem of hallucinations is. It seems to me it relies mostly on discounting just how much we've already had to deal with this same problem in humans over the millenia. The problem of proliferation of bad information might be getting worse, but this isn't native to generative AI. The entire informational ecosystem has to deal with this. GPTs compound the issue, but as…

If there's an upside here, it's that humans will now be forced to refine their BS detectors. In doing so the species would improve critical thinking skills which can be applied to all information regardless of source. Which, I agree, was often BS to begin with. But in theory would be more difficult to skirt on by without notice if humanity upgraded their critical thinking.

Unfortunately that didn't happen with other innovations of misinformation at scale (such as social media) so I'm not so optimistic.

Re: Generative AI could make search harder to trust

#127
I actually experienced this the other day. Bought the new Baldur's Gate and was wondering what items to keep or sell (don't judge me, I'm a pack rat in games!)

I had found some silver ingots. The top search result for "bg3 silver ingot" is a content farm article that very confidently claims you can use them at a workbench in Act 3 to upgrade your weapons.

Except this is a complete fabrication: silver ingots exist only to sell, and there is no workbench. There is no mechanic (short of mods) that allows you to change a weapon's stats.

I'm pretty sure an LLM "helped" write the article because it's a lot of trouble to go through just to be straight up wrong - if you're a low effort content farm, why in the world would you go through the trouble if fabricating an entire game mechanic instead of taking the low effort "They exist only to be sold" road?

This experience has caused me to start checking the date of search results: if it's 2022 and before, at least it's written by a human. If it's 2023 and on, I dust off my 90's "everything on the World Wide Web is wrong" glasses.

Re: Generative AI could make search harder to trust

#128

Earlier quoted context omitted.

People are vastly everestimating how unique this problem of hallucinations is. It seems to me it relies mostly on discounting just how much we've already had to deal with this same problem in humans over the millenia. The problem of proliferation of bad information might be getting worse, but this isn't native to generative AI. The entire informational ecosystem has to deal with this. GPTs compound the issue, but as…

Uh, no. Generative AI does not have a standard of truth. It generates text according to the probabilities given what it learned from its data set. There is no systematic modelling of a world that it can test its assertions against. What are called hallucinations are a fundamental property of the approach. The hallucinations are probable -- just not true. The model has succeeded but the assertions are false. This has…

[deleted]

Re: Generative AI could make search harder to trust

#129

Earlier quoted context omitted.

People are vastly everestimating how unique this problem of hallucinations is. It seems to me it relies mostly on discounting just how much we've already had to deal with this same problem in humans over the millenia. The problem of proliferation of bad information might be getting worse, but this isn't native to generative AI. The entire informational ecosystem has to deal with this. GPTs compound the issue, but as…

Uh, no. Generative AI does not have a standard of truth. It generates text according to the probabilities given what it learned from its data set. There is no systematic modelling of a world that it can test its assertions against. What are called hallucinations are a fundamental property of the approach. The hallucinations are probable -- just not true. The model has succeeded but the assertions are false. This has…

A standard of truth is doing a lot of heavy lifting. Myself I would say we assemble information based on the limitations of the body we exist in. For example gravity is important to us because if we disobey its laws we may very well die.

Embodiment and multi modal AI will likely provide such filters or limits to AI in which it can derive truth.

Re: Generative AI could make search harder to trust

#130
post #64

The SEO garbage has been poisoning the search for years. Even before the chatbots it got to the point when most top results are crap. The LLM's can surely make it much worse, though.

I was trying to do something with delegates in C# but couldn't remember what the various bits were called since it's been so long, so didn't have the magic words for Google to work. ChatGPT sorted me out with my vague question and one followup.

Ultimately, I would like to see more about the other side. If generative AI can make blog spam, then I think it can recognise blog spam. How far are we from implementing a reliable filter of useless spam sites from search results? I don't expect Google has a monetary incentive for this, but maybe someone else does. But from my story above maybe "search" is a thing of the past and for functional queries, we will just talk to the gatekeeper. Reading the actual words written will only be for leisure.

Post reply on HN