Live data from Hacker News

How I Made Google's "Web" View My Default Search

tedium.co

61–70 of 154 posts

Re: How I Made Google's "Web" View My Default Search

#61
post #51

Earlier quoted context omitted.

It's your ranking system, so it's correct by definition (for you), but I've got to point out you're not considering the quality of the text. In your model, an empty page is perfect. Cynically, maybe you're on to something...

> It's your ranking system, so it's correct by definition (for you), but I've got to point out you're not considering the quality of the text. Yes. This is by design! [1] > In your model, an empty page is perfect. Yes. What's the problem with that? An empty page won't match any search terms, will it? [1] It's impossible to quickly and cheaply determine the quality of the text, in an age where it's cheaper for a blog-…

You might be onto an interesting new view of the web here, if you've ever seen https://www.builtwith.com, something similar that also processes script tags and downweights based on what the script is would also add to this effect.

At the very least I'd be interested in looking at such a site and seeing how it differs from https://www.millionshort.com =)...

Re: How I Made Google's "Web" View My Default Search

#62
I don't mind the info boxes,... I get mad at google when I input two words into the search field, press enter, and the first few results don't include one of the words (50% of the search! .. even include the "show only links which include..."), and then uses synonym results for the second word, which gives out totally wrong results.

Re: How I Made Google's "Web" View My Default Search

#63

I want to make my own search engine, one day, with my own crawler. There is an SEO-proof way to determine what the ranking of a site should be - penalise it for each advertisement, penalise further for delivering different content to the crawler[1], allow logged-in users to down-rank a site, etc. Basically, a site starts off with a perfect score, then gets penalised for each violation, for each dark-pattern, for each…

I wish you luck and I hope you succeed, but you make it sound much much easier than what it would be.

First of all, you're going to drown in hardware costs, if you run your own hardware. If you run on AWS, you will be the largest AWS customer. 2 years ago, when Google was still displaying result counts, I got 1.3 billions results for "sushi"[1]. This means that if you use a reverse index to lookup your results, the "sushi" entry will be ~19GiB large, assuming you use UUIDs. If you think 90% of this is spam, and you only index the non spam (detecting spam/seo is far from trivial, but let's say you figure it out), you still need ~2GiB just for mapping "sushi". With 755,865 words in the English dictionary, according to wikipedia[2], you'll need ~1.5 PiB (yes, pebi/peta, 1,536 TiB) just to store relationships for English pages. This is assuming you don't support other languages, you discard 90% of pages, and you don't cache the content of pages for re-indexing.

In addition to this, you also need to store the meta-data for each pages (vote counts from your voting system, whether it's serving different content, etc...). The order of magnitude has to be in the O(100TiB) from my conservative gut feeling. (still assuming you discard 90% of the web, and I'll assume you aggregate the metadata on the domain, not on the individual pages)

The second challenge is your ranking. Now that you've become the dominant search engine with your awesome ranking system, you will become the main target for swaths of motivated click-farms which are exploiting workers from low income countries. They will be trying to register accounts, vote and game your ranking. You can most likely detect this behaviour, but their behaviour will be very similar to a significant portion of your real users. So you'll be fishing in a pond with a rocket launcher, and some of your legitimate users will be collateral victims. Otherwise, you'll spend most of your time playing a cat-and-mouse game with the SEO spammers instead of improving your search engine and fixing bugs.

I'm also falling in the trap "i could rewrite that in a weekend" sometimes, but for a search engine, I would love to see decent competition, but it's near impossible.

[1] https://news.ycombinator.com/item?id=30925402

[2] https://en.wikipedia.org/w/index.php?title=List_of_dictionar...

Re: How I Made Google's "Web" View My Default Search

#64

I want to make my own search engine, one day, with my own crawler. There is an SEO-proof way to determine what the ranking of a site should be - penalise it for each advertisement, penalise further for delivering different content to the crawler[1], allow logged-in users to down-rank a site, etc. Basically, a site starts off with a perfect score, then gets penalised for each violation, for each dark-pattern, for each…

> There is an SEO-proof way to determine what the ranking of a site should be

This comment deserves a "my sweet summer child". Everything can be gamed; if you don't see how, then you really shouldn't be suggesting ideas.

How do you define an advertisement? How do you prevent downvote spam (if captchas worked, there wouldn't be bots/astroturfers on the Internet)? How will you fund your search engine?

Re: How I Made Google's "Web" View My Default Search

#65

I want to make my own search engine, one day, with my own crawler. There is an SEO-proof way to determine what the ranking of a site should be - penalise it for each advertisement, penalise further for delivering different content to the crawler[1], allow logged-in users to down-rank a site, etc. Basically, a site starts off with a perfect score, then gets penalised for each violation, for each dark-pattern, for each…

> There is an SEO-proof way to determine what the ranking of a site should be This comment deserves a "my sweet summer child". Everything can be gamed; if you don't see how, then you really shouldn't be suggesting ideas. How do you define an advertisement? How do you prevent downvote spam (if captchas worked, there wouldn't be bots/astroturfers on the Internet)? How will you fund your search engine?

> How do you define an advertisement?

Headless browser with uBlock Origin. Check programmatically how many ads were blocked on the page.

Re: How I Made Google's "Web" View My Default Search

#66
post #26

Honest question - why still use Google at all? Whenever DDG became good enough (seven? ten? years ago?) I used it exclusively. Lately I use a mix and have moved on from DDG, but I still never went back to Google and don't understand how people can tolerate it; I find the results really bad.

I've using ddg for years but I find myself unconsciously ending up appending "!g" to almost all my searches. DDG results are just really not good, especially if I'm looking up something local.

I think you're right about local because small business owners have an incentive to optimise for Google local results in a way they don't for DDG or any other search engine.

Generally though, I find DDG good enough for everyday use. However, a friend of mine once told me she looked up a breed of bird on google for more information and I could barely contain my laughter when I suggested she should have used Duck Duck Go instead. The strange look I got in return reminded me that, for 99.9% of people, Google is 'Search' and that's the end of it.

Re: How I Made Google's "Web" View My Default Search

#67

Honest question - why still use Google at all? Whenever DDG became good enough (seven? ten? years ago?) I used it exclusively. Lately I use a mix and have moved on from DDG, but I still never went back to Google and don't understand how people can tolerate it; I find the results really bad.

For many of my searches DDG seems to completely ignore one or more of my keywords, usually giving me something more popular but less relevant. Almost every time when I try the same search on google it works.

I get the opposite effects on google. As an example, I tried to search for the documentation on watcom’s wasm assembler. It kept giving me results for web assembly no matter how I modified the query nor the Boolean search operators. To me it’s a symptom of search engines trying too hard to be smart and predictable.

Re: How I Made Google's "Web" View My Default Search

#68

Earlier quoted context omitted.

> There is an SEO-proof way to determine what the ranking of a site should be This comment deserves a "my sweet summer child". Everything can be gamed; if you don't see how, then you really shouldn't be suggesting ideas. How do you define an advertisement? How do you prevent downvote spam (if captchas worked, there wouldn't be bots/astroturfers on the Internet)? How will you fund your search engine?

> How do you define an advertisement? Headless browser with uBlock Origin. Check programmatically how many ads were blocked on the page.

But many ads and crap today are actually not traditional ads. Don’t they count? If not, then gpt generated text ads or whatever will become more prevalent

Re: How I Made Google's "Web" View My Default Search

#69
post #63

I want to make my own search engine, one day, with my own crawler. There is an SEO-proof way to determine what the ranking of a site should be - penalise it for each advertisement, penalise further for delivering different content to the crawler[1], allow logged-in users to down-rank a site, etc. Basically, a site starts off with a perfect score, then gets penalised for each violation, for each dark-pattern, for each…

I wish you luck and I hope you succeed, but you make it sound much much easier than what it would be. First of all, you're going to drown in hardware costs, if you run your own hardware. If you run on AWS, you will be the largest AWS customer. 2 years ago, when Google was still displaying result counts, I got 1.3 billions results for "sushi"[1]. This means that if you use a reverse index to lookup your results, the "…

Absolutely spot on. Additionally, it's worth mentioning that a lot of content is now locked behind a few major platforms (eg. Facebook, LinkedIn, Medium, YouTube, etc.) or CDNs like Cloudflare, which often block crawling from non-Google IPs or well-known search engines.

While the other costs mentioned here can be optimized with current hardware prices and a good database, anti-crawling measures necessitate thousands of IPs/proxies, making the process even more challenging and costly.

Re: How I Made Google's "Web" View My Default Search

#70
post #28

Earlier quoted context omitted.

I had to do this a while back and it's nearly made me punch my monitor. The lowest circle of hell for the person/people who decided putting this basic functionality behind a fucking arcane flag. Few things make me more mad than having my finite life wasted hunting down solutions to problems that have no reason to exist.

Apparently its behind a flag because the feature is not completely ready yet. However I was unable to find what work is still needed for it to be considered ready: https://bugzilla.mozilla.org/show_bug.cgi?id=1106626#c18

Not ready my ass, this was a thing before deciding to move away from XUL.
Post reply on HN