Live data from Hacker News

Generative AI could make search harder to trust

wired.com

141–150 of 255 posts

Re: Generative AI could make search harder to trust

#141

Earlier quoted context omitted.

Thats specific to recipes because they can’t be copyrighted

I had a look and definitely learned something today so #til. Also, note to self to collect my favourite recipes in markdown files from now on.

#til! https://www.justtherecipe.com/

Re: Generative AI could make search harder to trust

#143

I wonder if there will be a human information/knowledge equivalent of low-background steel (pre-WWII/nukes). Data from before a certain point won't be 'contaminated' with LLM stuff, but it'll be everywhere after that. https://en.wikipedia.org/wiki/Low-background_steel

I suspect in the coming years the Wayback Machine at archive.org will become ever more important - always assuming it's not lost as collateral damage in their copyright battles. Indexing that dataset and making it searchable would massively increase its value. My inner conspiracy theorist can't help wonder if the continued reduction in search usefulness isn't part of an ongoing deliberate disempowerment of everyday p…

It's a consequence of running search as a for profit system. When you're optimizing for revenue, the system priorities are different than if you were optimizing for user experience or true knowledge. Arguably that means that search should be a public utility but then you have to be able to trust your government with your searches and privacy.

Re: Generative AI could make search harder to trust

#144

Earlier quoted context omitted.

I suspect in the coming years the Wayback Machine at archive.org will become ever more important - always assuming it's not lost as collateral damage in their copyright battles. Indexing that dataset and making it searchable would massively increase its value. My inner conspiracy theorist can't help wonder if the continued reduction in search usefulness isn't part of an ongoing deliberate disempowerment of everyday p…

Generative Wayback Machine

Honestly, I would love to see a LLM based plugin that would take a website, remove all the tracking garbage and filler, and just give me back a no-frills static html+CSS site that looks like it was made in 1995.

"Oh, you want a guide to writing your own loss function for Tensorflow? Here's an FAQ that could have existed on comp.lang.python3.tensorflow"

Re: Generative AI could make search harder to trust

#145

I wonder if there will be a human information/knowledge equivalent of low-background steel (pre-WWII/nukes). Data from before a certain point won't be 'contaminated' with LLM stuff, but it'll be everywhere after that. https://en.wikipedia.org/wiki/Low-background_steel

Those simple web 1.0 sites made by college professors are a gold-standard in my book. I always enjoy finding them in search results. Although they are becoming increasingly rare.

Paul's Notes remains the absolute best textbook Calculus and intro to DiffEQ! Not sure if it's been updated since 2005, but I mean, it's not like they're discovering new Calc II methods! Being that it's not a $200 textbook rehashing the same stuff as the last 10 editions, it's easily one of my favorite websites

Re: Generative AI could make search harder to trust

#146

Earlier quoted context omitted.

Can't prove it, but it seems to me like black text on white background sites from the past are poorly ranked compared to sites with "modern" layouts.

Yes. I love black text on white background. A rare find these days. Browsing today is like: “You ask for a spaghetti recipe and the page tell you the whole history of civilization.”

Don't forget how Google will now drop search terms if it thinks you mean something else, or add unrelated synonyms to your search (presumably to "help" folks who aren't good at writing queries)

Re: Generative AI could make search harder to trust

#147

Earlier quoted context omitted.

> If it's 2023 and on, I dust off my 90's "everything on the World Wide Web is wrong" glasses. because misinformation written by humans didn’t exist before LLMs?

Back in the 90s when the Internet became a thing, it was common knowledge that because normal people made websites, that you should take things with a grain of salt. There was a bit of an overreaction to this, as the general feeling at the time was to trust nothing on the Internet. In the 00s and 10s, the quality of discoverable content improved: reddit and stackechange had experts (at a higher rate than the rest of…

Reddit never had experts. Maybe for half a minute. It became an echo chamber fast: fake internet points to be gained for saying what got upvoted last week, or to be lost for saying anything different.

Re: Generative AI could make search harder to trust

#148

Earlier quoted context omitted.

Back in the 90s when the Internet became a thing, it was common knowledge that because normal people made websites, that you should take things with a grain of salt. There was a bit of an overreaction to this, as the general feeling at the time was to trust nothing on the Internet. In the 00s and 10s, the quality of discoverable content improved: reddit and stackechange had experts (at a higher rate than the rest of…

Reddit never had experts. Maybe for half a minute. It became an echo chamber fast: fake internet points to be gained for saying what got upvoted last week, or to be lost for saying anything different.

I wanna politely disagree on that. I've found it to frequently be a resource on par with Stackechange.

Re: Generative AI could make search harder to trust

#149

I wonder if there will be a human information/knowledge equivalent of low-background steel (pre-WWII/nukes). Data from before a certain point won't be 'contaminated' with LLM stuff, but it'll be everywhere after that. https://en.wikipedia.org/wiki/Low-background_steel

https://en.wikipedia.org/wiki/World_Brain

But who's gonna pay for it?

Re: Generative AI could make search harder to trust

#150
post #136

I think it'd be kind of neat in a backward way if we went back to the 'specialized encyclopedia' days of the 90s. Web directories , 'Who's Who in Engineering' type lists, etc. It's a step back from universal search engines being able to find stuff, but it's a step forward with regards to curation and quality of results; so i'm not sure if it's entirely a downgrade. The early 90s 'website phonebook' type encyclopedias…

It's already kind of like that for me, in that almost all of my searches fall into these categories:

- Wikipedia

- Online documentation for whatever language/framework/tool I'm using

- Stack Overflow / Stack Exchange for most technical questions

- Reddit if SO/SE doesn't work, and for opinionated questions (e.g. r/BuyItForLife)

- Hacker news for software recommendations and technical opinionated questions

- Arxiv or the ACM library if it's a research paper (99% of the time, whenever I google something niche the only relevant results are papers)

- Other sites like caniuse.com, university sites for health and nutritional info, old-style forums for specific software

For these searches I'm just using Google to bring me to the specific site I want, because it's faster than using the site's own search functionality. Then there are the times I literally just type in the website instead of the URL bar (e.g. "instacart"), or when I use Google maps, images, or reviews.

I'm always wary when Google returns an unfamiliar site because I'm skeptical of the results. ~70% of the time it's some blogspam which is at best accurate but overly wordy, and at worst inaccurate; sometimes it's a blog from some random individual who for whatever reason went into a deep dive trying to understand what I'm searching for, that actually turns out to be useful; the rest, idk.

Post reply on HN