Live data from Hacker News

Google's Results Are Infested, Open AI Is Using Their Playbook from the 2000s

chuckwnelson.com

321–330 of 504 posts

Re: Google's Results Are Infested, Open AI Is Using Their Playbook from the 2000s

#321

Earlier quoted context omitted.

What I’m getting at is simple, no one is going to find a random persons obscure blog where they are trying to build a “brand” or be a “thought leader” that is not on the first page of search results. I subscribe to Ben Thompson’s writing and make it habit to go to a few other websites because they have earned my trust. The only method that most people have ever had of gaining traction is via word of mouth and not thr…

I don't know how old you are, or whether you ever really knew the web in the prior era that we're talking about. Forgive me if I'm making flawed guesses about where you're coming from. Back in the day, if I wanted the answer to some specific question about, say, restaurants in Chicago, I'd search for it on Google. Even if I didn't know enough about the topic to recognize the highest quality sites, it was okay, becaus…

I’m old enough that my first paid project was making modifications to a home grown Gopher server built using XCMDs for HyperCard.

My first post was on Usenet in 1994 using the “nn” newsreader

The web has gotten much larger than when it didn’t exist when I started.

But web rings on GeoCities weren’t exactly places to do “high quality research”. You still had to go to trusted sites you knew about or start at Wikipedia and go to citations.

Before two years ago I would go to Yelp. Now I use the paid version of ChatGPT that searches the internet and returns sources with links

https://imgur.com/a/hZwrjJS

Re: Google's Results Are Infested, Open AI Is Using Their Playbook from the 2000s

#322

Earlier quoted context omitted.

But they can never be. RAG gets you somewhere, but it’s still a pile of RNGs under a trenchcoat.

>> ideally

It’s just not possible. You can do a lot with nondeterministic systems, they have value - but oranges and apples. They need to coexist.

Re: Google's Results Are Infested, Open AI Is Using Their Playbook from the 2000s

#323
post #211

Don't bet on AI staying clean. A lot of HN readers conceptualize the forces attacking the integrity of the search results as just some isolated people taking occasional potshots, and then maybe slinking away if their trick gets blocked. It is probably a lot more accurate to visualize the SEO industry as a Dark Google. Roughly as well resourced, with many smart people working on it full time, day in, day out, with inf…

"dark Google" seems like the title of a blog post I would find on HN! This is intended as a compliment, in case not clear... Add some important facts and figures (what is the revenue of dark Google, who and how many are they employing) and write it up!

Re: Google's Results Are Infested, Open AI Is Using Their Playbook from the 2000s

#324

Earlier quoted context omitted.

It invents citations too, constantly. You could look up the things it cites, although at that point, what are you actually gaining? And I’m not saying this makes them useless: I pay for Claude and am a reasonably happy customer, despite the occasional bullshit. But none of that is relevant to my point that the bots get held to a different standard than Google search and I don’t see an easy way for Google to deal with…

Do you pay for ChatGPT? The paid version of ChatGPT has had a web search tool for ages. It will search the web and give you live links.

ChatGPT has had web search for exactly 58 days. I guess our definitions of 'ages' differ by several orders of magnitude.

Re: Google's Results Are Infested, Open AI Is Using Their Playbook from the 2000s

#325

I wouldn't be concerned about trusting the results of ChatGPT if it also were providing links to the sources it had cited or used as a reference in its answers. Unfortunately, it doesn't, and so I can't verify them. Not sure if it's an actual limitation of current LLM or rather they're intentionally filtering out the sources.

> I wouldn't be concerned about trusting the results of ChatGPT if it also were providing links to the sources it had cited or used as a reference in its answers.

This will only be a problem for a few more years. Soon every article, paper, and website will be generated by a LLM. Verifying the output of ChatGPT by referencing other LLM generated source material will be a pointless exercise.

Re: Google's Results Are Infested, Open AI Is Using Their Playbook from the 2000s

#326
post #314

Earlier quoted context omitted.

I very much agree this is effectivity a 'honeymoon' period. Expect the SEO collective to shift focus on AI if the search approach becomes profitable in a few years. That said, given an "AI search" is estimated to be at least ten times [0] as expensive per query than traditional search, I hope you like ads. For those hoping to see that cost to go down, training costs for improved models have instead been going up . [1…

> I very much agree this is effectivity a 'honeymoon' period. At this point I'd be much more interested to hear which "unicorn" tech company did not have such a honeymoon period which it later turned away from. This should really be the default, expected behaviour at this point.

> At this point I'd be much more interested to hear which "unicorn" tech company did not have such a honeymoon period which it later turned away from

Doctolib in France (and Italy, Germany, Netherlands) is one such example. Founded in 2013 so decent life, still as good as in the beginning for both consumers (people booking healthcare appointments) and the customers (doctors paying to use it for their appointment management). And they're only getting better, with e.g. an AI assistant in beta to take notes during appointments.

Re: Google's Results Are Infested, Open AI Is Using Their Playbook from the 2000s

#327
post #41
post #7

> Enter 2024 with AI. The top 20% of search results are a wall of text from AI... I'll be the contrarian here and say I actually like Google's AI Overview? For the first time in a long time, I can search for an answer to a question and, instead of getting annoying ads and SEO-optimized uselessness, I actually get an answer. Google is finally useful again. That said, once Google screws with this and starts making sear…

The answers come from the same websites. They just get stripped of their traffic. As someone who puts a ton of work into writing accurate, helpful guides, it's devastating to have my work plundered like that. Once these monopolies have successfully established themselves, they will become indistinguishable from the ad-invested websites they replace. The only difference is that they will create no new information of t…

[dead]

Re: Google's Results Are Infested, Open AI Is Using Their Playbook from the 2000s

#328
> When Google came onto the scene, I credit its success to the tried and true paradigm that makes companies successful: simple and easy to use.

> Yahoo was dominant back then, and it tried to put everyone and everything in front of you. Then we learned about the paralysis of choice. Too many choices, the mental fatigue weighed in, and the product became difficult to use.

This nonsense again? I was around then, and I switched from Yahoo and AltaVista to Google despite its dumb name and stupid, childish logo because Google's results were hands-down better. Instead of a solely full-text search paradigm based only on keyword density, Google also ranked pages based on how many other pages linked to them, the so-called "PageRank" algorithm.

This worked much, much better, and was much harder (for a while) to game. Before Google, it was common when searching to find pages that gamed the search engines by stuffing their keyword tags with SEO crap or putting it in giant footer sections in a tiny font the same color as the background (to render it invisible). Google's PageRank wasn't fooled by this.

Also most of the major search engines adopted similarly minimalist UIs, and it did zero to stop the bleeding. They all lost to google. (AltaVista, the pre-Google Google, was still useful for a while for some specialty searching, like for anonymous FTP servers, and I wonder if DEC had never gone under or if Compaq had spun off AltaVista, maybe history would be different.)

EDIT: I just realized the article doesn't even mention AltaVista. Unbelievable.

Re: Google's Results Are Infested, Open AI Is Using Their Playbook from the 2000s

#329
post #293

Earlier quoted context omitted.

Funny thing is you can train a small BERT model to detect queries that are in categories that aren’t ready for AI “answers” with like .00000001% of the energy of an LLM.

That's (obviously) a bit of an exaggeration. BERT is just another transformer architecture. Cut down from ~100 layers to 1, ~1k dimensions to ~10, and ~10k tokens to 100, and you're only 1e6 faster / more efficient, still a factor of 10k greater than your estimate and also too small to handle the detection you're describing with any reasonable degree of accuracy.

I literally have DistilBERT models that can do this exact task in ~14ms on an NVIDIA A6000. I don’t know the precise performance per watt, but it’s really fucking low.

I use LLM to help with training data as they are great at zero shot, but after the training corpora is built a small, well trained, model will smoke an LLM in classification accuracy and are way faster - which means you can get scale and low carbon cost.

In my personal opinion there is a moral imperative to use the most efficient models possible at every step in a system design. LLM are one type of architecture and while they do a lot well, you can use a variety of energy efficient techniques to do discrete tasks much better.

Re: Google's Results Are Infested, Open AI Is Using Their Playbook from the 2000s

#330
post #261

Earlier quoted context omitted.

Hospitals are increasingly owned by insurance companies. The customers are not doctors but shareholders. That is why a cure is seen as a threat.

That doesn’t make sense - a sick patient costs the shareholders money.

Only if you treat the patient. Cures cost money to administer. Better to just deny cover in the first place.
Post reply on HN