Live data from Hacker News

Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

danluu.com

391–400 of 493 posts

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#391

The intro query "youtube downloader" already showed me relevant results (some website where you paste an URL and bam download). I think there's a big tech bias in the whole post (how relevant is a mastodon poll, for real). Not saying the current landscape doesn't suck with ads everywhere and incentives to not give exactly relevant results at times, but I think google is pretty good still.

Which web site did you use to successfully download a youtube video, and which youtube video did you downloadl?

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#392

Weird article. Basically, the author thinks that anything that is not yt-dlp is a bad search result, which is pretty insane. Like, for me at least, I already know yt-dlp exists. When I search "youtube downloader", it's exactly because I want an online-website page to download youtube videos.

The author would probably accept any result that helps them download youtube videos. Did you find any and successfully use it to download a youtube video? Could you provide a link to the one you used?

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#393

Earlier quoted context omitted.

I use Kagi because I'm trying to remove Google from my life, but their text search is worse than Google in my experience, and the image search is abysmal. I'm wondering how long I can keep this up. I already revert to Google for image search, and am finding myself using either Google or ChatGPT over Kagi more and more for text as well.

Kagi had a pretty substantial image search update just few days ago [1]. Do you still the issues with it? [1] https://kagi.com/changelog#2793

Good info - will experiment!

It's already performing better on a (n=1) test I tried.

"Talos Principle 2". (Video game sequel) Previously (~5 days ago), Google returned various screenshots etc from the game `The Talos Principle 2`. Kagi returned mostly results from `The Talos Principle (1)`. Now the latest Kagi results are a mix, mostly from 2. So, it does look like it fixed this query.

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#394

Blah blah blah. Could you lay this article out any worse? What are the queries you used to test? I want to try them too. Buried in here somewhere. Using an adblocker is not expert anything. That you've defined your own opinion for what some of the results should be blows the thing up. Searching youtube downloader, many people would be fine with some of the ad covered but totally functional sites that pop up on Google…

Which web site did you use to successfully download a youtube video, and which youtube video did you download?

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#395
post #277

Earlier quoted context omitted.

> I really don't understand why anyone writing articles about ChatGPT uses 3.5. Because that’s what most people have access to. It’s absolutely worthless to most readers to talk about something they’ll never pay for and it’s not the job of random third-parties to incentivise others to send money to OpenAI. What I really don’t understand is why anyone gets so hung up about it and blames the writer. If you’re bothered…

> Because that’s what most people have access to. I’d agree with this rationale if the author clearly communicated their choice of model and the consequences of that choice upfront. In this post the table of results and the text of the post itself simply reads “ChatGPT” with no mention of 3.5 until the middle of a paragraph of text in the appendix. > It’s absolutely worthless to most readers to talk about something t…

> I’d agree with this rationale if the author clearly communicated their choice of model and the consequences of that choice upfront. (…) with no mention of 3.5 until the middle of a paragraph of text in the appendix.

You’re moving the goalposts. You went from criticising anyone using 3.5 and writing about it to saying it would’ve been OK if they had mentioned it where you think it’s acceptable. It’s debatable if the information needed to be more prominent; it is not debatable it is present.

> If you simply say “ChatGPT” it’s reasonable to infer that you’re evaluating the best possible version of “ChatGPT”, not the worst.

Alternatively, it you simply say “ChatGPT” it’s reasonable to infer that you’re evaluating the version most people have access to and can “play along” with the author.

> If using GPT4 vs 3.5 would create results so distinct from one another that it would serve to incentivize people to give money to OpenAI

Those are your words, not mine. I argued for the exact opposite.

> Again, if they’re making money off their readers it’s their job to provide them with an accurate representation of the tech.

I agree they should strive to provide accurate information. But I disagree that being paid has anything to do with it, and that their representation of the tech was inaccurate. Incomplete, maybe.

> Regardless, if this “excessive fawning” is truly unwarranted, this would again undermine your statement that using GPT4 would “incentivize others to send money to OpenAI”.

Again, I did not argue that, I argued the opposite. What I meant is that even if you believe that to be true, that still doesn’t mean random third-parties would have any obligation to do it.

> I’ll highlight what another commenter replied to you.

That comment has a reply, by another person, to which I didn’t feel the need to add.

> It seems in this case you’re holding ChatGPT to an arbitrary standard, not to mention one that the majority of humanity, including many of its brightest members, would fail to meet.

Machines and humans are not the same, not judged the same, don’t work the same, are not interpreted the same. Let’s please stop pretending there’s an equivalence.

Here’s a simple example: If someone tells you they can multiply any two numbers in their head and you give them 324543 and 976985, when they reply “317073642855” you’ll take out a calculator to confirm. If you had done the calculation first on a computer, you wouldn’t turn to the nearest human for them to confirm it in their head.

The problem with ChatGPT being wrong and misleading isn’t the information itself, but that people are taking it as correct because that’s what they’re used to and expect from machines. In addition, you don’t know when an answer is bullshit or not. With a human, not only can you catch clues regarding reliability of the information, you learn which human to trust with each information.

Everyone’s standard for ChatGPT, be it absolute omniscience, utter failure, or anything in between, is arbitrary. Comparing it to “the majority of humanity, including many of its brightest members” is certainly not an objective measurable standard.

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#396

The issue with traditional search engines is that keyword-first algorithms are extremely gameable. Try https://search.metaphor.systems - it's fully neural embeddings-based search. No keywords, only an embedding of what the actual content of a webpage is. So in the mentioned example of searching for Youtube downloaders, with Metaphor you'll get only Youtube downloaders ( https://search.metaphor.systems/search?q=This%2…

How do you deal with dynamically/contextually generated content? And how about paywalls and login-required content?

Do our best at getting the right content.

For paywalls/login - we play pretty straight, always obey robots.txt, etc.

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#397

Earlier quoted context omitted.

Yup, definitely still gameable but if the model learns what high quality content is like and what high quality webpages there are (which it does), then the only way to game would be to be great :) For your search - I would recommend turning autoprompt off and searching something like "Here is a great summary of the best computer mice to use:". Our embeddings model is trained on how links are talked about on the Inter…

> Our embeddings model is trained on how links are talked about on the Internet, if that helps with querying. So you have to query like how someone would refer to a link before sharing it So it's not high quality web pages but web pages that people talk about a lot which is expected since no one has an oracle that says what high quality is. The embeddings are merely a proxy and generalization for "how links are talke…

That's true. Although should be much harder

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#398

What always confuses me about the „search has gotten so bad“ mentality is that it is often based on anecdotal evidence at best, and anecdotal recollection at worst. Like, sure, I have the impression that search got worse over the last years, but .. has it really? How could you tell? And, honestly, this should be a verifiable claim; you can just try the top N search terms from Google trends or whatever and see how the…

To do this you would need to have a comprehensive definition of "quality", and that's anything but easy, and it will be at least partly subjective. It's also hard to include omissions in your definition of "quality" (and again, what should or should not be omitted is subjective as well).

For example, let's say I search for "Gaza"; on one extreme end some engines might only focus on recent events, whereas others may ignore recent events and includes only general information. Is one higher "quality" than the other? Not really – it depends what you're looking for innit?

All you can really do is make a subjective list of things you find important and rate things according to that, and this is basically just the same thing as an anecdotal account but with extra steps.

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#399

Can someone tell me why Bing, and thus DDG, has switched to prioritizing local results? I'll search the most inane things, like lyrics to a song, and get results for local businesses containing maybe one word in common. It's most frustrating with phone numbers. I picked up the habit of searching the random numbers that called me, to try and find out if they were possibly important. I used to get a bunch of spam sites…

> Bing, and thus DDG, has switched to prioritizing local results From what I can tell this is an issue with the Bing API that DDG uses that the DDG folks have been unable to resolve. I've tried many identical queries between DDG and Bing and while Bing does occasionally return incorrect local results, the completely irrelevant local results that appear on almost every DDG search do not seem to happen with Bing itself…

Long time DDG user (>10 years) here, and it’s astounding to me that they haven’t prioritized making their own independent index to switch off Bing. I would have expected them to do it like 5 years ago, but there’s afaik no initiative to do so. It’s unfortunate and am now trying other engines like Brave search.

Re: Compare Google, Bing, Marginalia, Kagi, Mwmbl, and ChatGPT

#400

What always confuses me about the „search has gotten so bad“ mentality is that it is often based on anecdotal evidence at best, and anecdotal recollection at worst. Like, sure, I have the impression that search got worse over the last years, but .. has it really? How could you tell? And, honestly, this should be a verifiable claim; you can just try the top N search terms from Google trends or whatever and see how the…

> has it really? How could you tell?

Yes it has and for a certain class of queries it's not even open for debate, because Google themselves have stated they deliberately made it worse. And they really did, it's very noticeable.

This class of queries is for anything related to any perspective deemed "non authoritative". Try to find information that contradicts the US Government on medical questions, for example, and even when you know what page you're looking for you won't be able to find it except via the most specific forcing e.g. exact quoted substrings.

Likewise, try finding stories that are mostly covered by Breitbart on Google and you won't be able to. They suppress conservative news sites to stop them ranking.

15 years ago Google wasn't doing that. It would usually return what you were looking for regardless of topic. There are now many topics - which specifically is a secret - on which the result quality is deliberately trashed because they'd prefer to show you the wrong results in an attempt to change your mind about something, than the results you actually asked for.

Post reply on HN