Live data from Hacker News

Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

danluu.com

301–310 of 493 posts

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#301

What always confuses me about the „search has gotten so bad“ mentality is that it is often based on anecdotal evidence at best, and anecdotal recollection at worst. Like, sure, I have the impression that search got worse over the last years, but .. has it really? How could you tell? And, honestly, this should be a verifiable claim; you can just try the top N search terms from Google trends or whatever and see how the…

Internet Archive remembers. https://web.archive.org/web/*/google.com/search/%2A

Find a query of interest, see for yourself (and take a snapshot of the present state for posterity).

The api enables more powerful queries, https://web.archive.org/cdx/search/cdx?url=google.co.jp*&pag...

Also try other search engines and languages.

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#302
post #100

Can someone tell me why Bing, and thus DDG, has switched to prioritizing local results? I'll search the most inane things, like lyrics to a song, and get results for local businesses containing maybe one word in common. It's most frustrating with phone numbers. I picked up the habit of searching the random numbers that called me, to try and find out if they were possibly important. I used to get a bunch of spam sites…

Yeah, I've noticed this as well with DDG recently: even with the localised checkbox disabled it still prioritises them, which often is very frustrating as the results are then almost totally useless. However, more generally, I've personally found that DDG (and maybe Bing's then?) localised results are just really bad, and have been for the multiple years I've been using DDG and it's had this feature: I'm in New Zeala…

I habe the same experience from Germany. There's the slider but it's not doing mich.

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#303
post #207

I'm in the camp of those who think Google's results are still very good. I admit I use adblock (uBlock Origin) and won't even try to disable it. I understand the author's point of turning off their ad blocker "to get the non-expert browsing experience" but then they could make a different test with uBlock on for every query and see how it goes. It's also a bit inconsistent to expect results for downloading videos men…

[deleted]

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#304

I'm not able to reproduce the author's bad results in Kagi, at all. What I'm seeing when searching the same terms is fantastic in comparison. I don't know what went wrong there. In the Youtube Downloader search, NortonSafeWeb is nowhere to be found. I get a couple of legit downloader websites, and some articles from reputable tech newspapers on how to use them or command line tools. In the Adblock search, ublock Orig…

[deleted]

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#306

Earlier quoted context omitted.

I'm curious: what is the rationale for "in an incognito tab" being part of the test harness? It seems pretty arbitrary to me to disable one of the key features - in this case personalization - of the software being evaluated. Or is the evaluation not between "search engines" but rather "search engines without personalization"? If so, then this restriction does make sense. But that is not the evaluation that "normal u…

> I'm curious: what is the rationale for "in an incognito tab" being part of the test harness? It's the closest we can easily get to the 'average user experience'. Someone who has a long account/cookie history with Google has plausibly trained the site to return more relevant results through implicit user-curation of avoiding obvious-to-them SEO-spam on other queries. If we posit that every user eventually trains Goo…

> It's the closest we can easily get to the 'average user experience'.

Maybe it's the closest we can get (though I doubt it), but it definitely isn't close enough to tell us anything about the "average user experience".

The average user has been using google for years, without taking any steps to avoid personalization. An incognito session (on a browser / machine / network that is probably fingerprinted...) is pretty much the opposite of that typical usage pattern.

I recognize that just writing a blog post or comment on HN is not a research project so needs to do something quick, but I think it mostly invalidates the experiment. What would get closer would be to devise a few user personas and attempt to search and browse for awhile within those personas before trying the experiment. Or much better yet, put together a focus group comprised of real people within the personas you're interested in, and run the experiment using their real accounts.

> If we posit that every user eventually trains Google to avoid SEO spam

I don't think it's that, I think it's that every user trains it to return results more likely to improve the metric of "more likely to click one of the links", and I think that makes it more, not less, likely that they see what most of us here consider to be spam.

But I don't know! Maybe that's not what this experimental setup would show. But it would be a lot more enlightening than a setup using a fresh incognito window, which reflects the usage pattern of a proportion of search queries that is a tiny rounding error above zero.

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#307

What always confuses me about the „search has gotten so bad“ mentality is that it is often based on anecdotal evidence at best, and anecdotal recollection at worst. Like, sure, I have the impression that search got worse over the last years, but .. has it really? How could you tell? And, honestly, this should be a verifiable claim; you can just try the top N search terms from Google trends or whatever and see how the…

> Dan at least started to provide actual evidence and criteria by which he would score results, but even he only looked at 5 examples. Which really is a small sample size to make any general claims.

US NIST, in their annual TREC evaluation of search systems in the scientific/academic world, use sets of 25 or 50 queries (confusingly called "topics" in the jargon).

For each, a mandated data collection is searched by retired intelligence analysts to find (almost) all relevant result, which are represented by document ID in general search and by a regular expression that matches the relevant answer for question answering (when that was evaluated, 1998-2006).

Such an approach is expensive but has the advantage of being reusable.

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#308

What always confuses me about the „search has gotten so bad“ mentality is that it is often based on anecdotal evidence at best, and anecdotal recollection at worst. Like, sure, I have the impression that search got worse over the last years, but .. has it really? How could you tell? And, honestly, this should be a verifiable claim; you can just try the top N search terms from Google trends or whatever and see how the…

Dan approached the problem from a qualitative perspective. Perhaps if more people took this approach over quantitative maximalism we would actually have products that don’t drive us fucking insane.

All that matters is the overwhelming sentiment that search has gotten worse, not the same fucking spreadsheet that got us here in the first place!

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#309

Earlier quoted context omitted.

I also assume that Kagi uses some shady residential IPs proxies and similar tricks to scrap Google while DDG has access to the Bing API.

You can buy access to the Google Search API, which is what I assume Kagi does. Building your product on being able to circumvent some Google restrictions seems like a bad business move, if you can buy the same service for a reasonable price.

Where can I buy it?

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#310

Earlier quoted context omitted.

I think the point he's trying to make that the search results page from the mainstream search engines are a minefield of scams that a regular person would have difficulty navigating safely. If he was looking at relevance, yours would be a solid point, but since most of the emphasis is on harm, a smaller sample works. Like "we found used needles in 3 out of 5 playgrounds" doesn't typically garner requests for p-values…

I think this is a good illustration of my frustration with this discussion: I don't think search has gotten bad, I think the web has gotten bad. It's weird to even conceptualize it as a big graph of useful hypertext documents. That's just wikipedia. The broader web is this much noisier and dubious thing now. That's bad for google though! Their model is very much predicated on the web having a lot of signal that they…

On the one hand, I'm not sure the data corroborates that. If this is a web problem and not a search engine problem, then I'd expect every search engine to have the same pattern of scam results.

I'd also argue that finding relevant results among a sea of irrelevant results is the primary function of a search engine. This was as true in 1998 as it is today. In fact, it was Google's "killer feature", unlike Altavista and the likes it showed you far more relevant results.

Post reply on HN