Live data from Hacker News

How I Made Google's "Web" View My Default Search

tedium.co

101–110 of 154 posts

Re: How I Made Google's "Web" View My Default Search

#101

I don't mind the info boxes,... I get mad at google when I input two words into the search field, press enter, and the first few results don't include one of the words (50% of the search! .. even include the "show only links which include..."), and then uses synonym results for the second word, which gives out totally wrong results.

It infuriates me that exact match searches don't work anymore. I understand that for the average person/average query, the modern google results are better (with info boxes, AI type stuff going on), but please let me type in "I want this exact phrase, including punctuation!". I don't even care if it tells me there's nothing found and then makes some suggestions below that.

Google has restored 'verbatim' search, as a drop-down beneath "all results." If you'd prefer to add it to a bookmarklet, the URL parameter is tbs=li:1.

Re: How I Made Google's "Web" View My Default Search

#102
post #44

I want to make my own search engine, one day, with my own crawler. There is an SEO-proof way to determine what the ranking of a site should be - penalise it for each advertisement, penalise further for delivering different content to the crawler[1], allow logged-in users to down-rank a site, etc. Basically, a site starts off with a perfect score, then gets penalised for each violation, for each dark-pattern, for each…

> allow logged-in users to down-rank a site, etc. And then you need a huge anti-bot mechanism if the search engine will get any popularity. There is money-based motivation to affect the rankings.

There's also the issue of users trying to downrank any polarizing site because they disagree with it. To pick an example, think of the search term "vaccine" and how strongly people on both sides feel about it. Eventually, you'll probably end up with results for "vaccine" that have nothing whatsoever to do with vaccines because everything vaccine-related will end up downranked.

Re: How I Made Google's "Web" View My Default Search

#103
post #69
post #63

Earlier quoted context omitted.

I wish you luck and I hope you succeed, but you make it sound much much easier than what it would be. First of all, you're going to drown in hardware costs, if you run your own hardware. If you run on AWS, you will be the largest AWS customer. 2 years ago, when Google was still displaying result counts, I got 1.3 billions results for "sushi"[1]. This means that if you use a reverse index to lookup your results, the "…

Absolutely spot on. Additionally, it's worth mentioning that a lot of content is now locked behind a few major platforms (eg. Facebook, LinkedIn, Medium, YouTube, etc.) or CDNs like Cloudflare, which often block crawling from non-Google IPs or well-known search engines. While the other costs mentioned here can be optimized with current hardware prices and a good database, anti-crawling measures necessitate thousands…

Additionally, it's worth mentioning that a lot of content is now locked behind a few major platforms (eg. Facebook, LinkedIn, Medium, YouTube, etc.) or CDNs like Cloudflare, which often block crawling from non-Google IPs or well-known search engines.

I think this is fine. If I want to find something on one of those big sites I just go there directly. However if I want to search the web for a site I’ve never been to before then I’m stuck with the bad results of the current search offerings. It’s quite depressing!

Re: How I Made Google's "Web" View My Default Search

#104

I want to make my own search engine, one day, with my own crawler. There is an SEO-proof way to determine what the ranking of a site should be - penalise it for each advertisement, penalise further for delivering different content to the crawler[1], allow logged-in users to down-rank a site, etc. Basically, a site starts off with a perfect score, then gets penalised for each violation, for each dark-pattern, for each…

Wikipedia shows an ad asking for money.

So a hundred times a day, I open a new Wikipedia clone with no ads. I get perfect scores.

Once I detect your bot has crawled it, I put ads all over it until the traffic finally goes away.

No system is fool-proof.

Re: How I Made Google's "Web" View My Default Search

#105
I solved this Google problem by paying someone else for search. I have switched to kagi.com, a so-called, "paid, ad-free search engine".

I have not yet found one drawback to using it. Unlike previous attempts to switch from Google, I have not ever felt like I was getting substandard answers or thought of switching back. It's been months.

And, it is lovely. The lack of ads and of it having no manipulative motivation makes me very happy. It has other features that are nice.

Re: How I Made Google's "Web" View My Default Search

#106
post #82

Earlier quoted context omitted.

>but it appears to me that it dismisses all commercial content altogether, which is not what I want Isn't it? You want to downrank for ads and downrank for paywalls. How is commercial content supposed to be funded?

> >but it appears to me that it dismisses all commercial content altogether, which is not what I want > Isn't it? Of course not. Appearing lower in the results than non-monetised content is very different from not appearing in the results at all. > How is commercial content supposed to be funded? They'll find a way. After all, if more people turn to ChatGPT for queries than to search engines, those sites are under th…

>Appearing lower in the results than non-monetised content is very different from not appearing in the results at all.

I'd say it's only technically different unless there are only a handful of results. It probably might as well not appear in results if it's past the twentieth position. On Google, click through rate is already down to ~1% by the tenth result.

Re: How I Made Google's "Web" View My Default Search

#107

I solved this Google problem by paying someone else for search. I have switched to kagi.com, a so-called, "paid, ad-free search engine". I have not yet found one drawback to using it. Unlike previous attempts to switch from Google, I have not ever felt like I was getting substandard answers or thought of switching back. It's been months. And, it is lovely. The lack of ads and of it having no manipulative motivation m…

I have been using Brave Search for a few years now and it's fantastic as well. I don't feel the need to use Google anymore and even if I do for images, I just add a !g to the search query and it automatically redirects me to Google.

Re: How I Made Google's "Web" View My Default Search

#108
post #82

Earlier quoted context omitted.

>but it appears to me that it dismisses all commercial content altogether, which is not what I want Isn't it? You want to downrank for ads and downrank for paywalls. How is commercial content supposed to be funded?

> >but it appears to me that it dismisses all commercial content altogether, which is not what I want > Isn't it? Of course not. Appearing lower in the results than non-monetised content is very different from not appearing in the results at all. > How is commercial content supposed to be funded? They'll find a way. After all, if more people turn to ChatGPT for queries than to search engines, those sites are under th…

They'll find a way.

If a site doesn’t appear on a Google results page, it dies. If a site doesn’t appear on your results page, people switch back to Google to search for it.

As a new search engine you’re going to be faced with this sort of thing for a long time before users begin to trust you. If a user actually wants to find a commercial site but simply can’t, they’re not going to stick around for long.

Re: How I Made Google's "Web" View My Default Search

#109

I want to make my own search engine, one day, with my own crawler. There is an SEO-proof way to determine what the ranking of a site should be - penalise it for each advertisement, penalise further for delivering different content to the crawler[1], allow logged-in users to down-rank a site, etc. Basically, a site starts off with a perfect score, then gets penalised for each violation, for each dark-pattern, for each…

Wikipedia shows an ad asking for money. So a hundred times a day, I open a new Wikipedia clone with no ads. I get perfect scores. Once I detect your bot has crawled it, I put ads all over it until the traffic finally goes away. No system is fool-proof.

But there's a difference between an embedded ad and a native ad. You can allow websites to just add an image or a section of an ad that does not hit a third-party endpoint.

Native ads are not as intrusive and if they become intrusive, like a popup, you can penalize that.

It sounds like a sound system.

Re: How I Made Google's "Web" View My Default Search

#110
post #16

Earlier quoted context omitted.

> why get their site to the top of the rankings if they can't put any advertisements on it? You would need to penalize any website that accepts money from users. There are a lot of SaaS companies out there writing blog spam to promote subscription products.

> You would need to penalize any website that accepts money from users. There are a lot of SaaS companies out there writing blog spam to promote subscription products. I don't see that as a problem, because blog-spam without advertisements is going to get down-ranked until it is exactly where it needs to be. If a single site dominates too much, then a hard cap on how how often a specific site can appear in the result…

Specific exceptions can be made for wikipedia, or similar.

Once you go down the dark path of white-listing sites, you’ve admitted defeat. How are you going to find that upstart Reddit competitor or federated Wikipedia competitor?

Post reply on HN