Live data from Hacker News

To Break Google’s Monopoly on Search, Make Its Index Public

bloomberg.com

521–530 of 630 posts

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#521
I don't think this would do anything. Indexing the web is compartively easy next to querying that index to find results relevant to the users query.

On top of this, Google search is as good as it is thanks in no small part to business practices that make Google as a company unsavory. Tracking your browsing habits, reading your emails, and watching your geographical location all feed the search algorithm, providing more seemingly prescient search results. For people who are OK with this, I don't see how an alternative can show up and compete with the years of data Google has already collected on a given user.

Toppling the Google monopoly is going to be more about a cultural shift than just giving competitors the data and algorithms they need to become viable. Users need to be made more aware of how Google uses their data, then alternatives can pop up marketing themselves on how they do (or don't) leverage user data. I'm skeptical that enough users will actually care to make this difference, at least in the short term.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#522

Ex-Google-Search engineer here, having also done some projects since leaving that involve data-mining publicly-available web documents. This proposal won't do very much. Indexing is the (relatively) easy part of building a search engine. CommonCrawl already indexes the top 3B+ pages on the web and makes it freely available on AWS. It costs about $50 to grep over it, $800 or so to run a moderately complex Hadoop job.…

If there were viable alternatives, people would shift over time. If I type in “ Pentagon” on Google, the first link is LinkedIn. DuckDuckGo doesn’t even list it at all. There’s countless examples where DuckDuckGo just can’t find basic information. DDG is just unreliable beyond it’s silly name.

I find DDG has pretty acceptable or even good results most of the time.

The real power is in the "bangs", though; you can use the `!` to immediately jump to the first search result without seeing a search page, or use `!g` to switch to Google for this particular query, among others. It enables a sort of power-user usage that one wouldn't get with Google.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#523
post #361

Earlier quoted context omitted.

This ^ times a 1000. Google simply has the best search product. They invest in it like crazy. I’ve tried bing multiple times. It’s slow, it spams msn ads in your face on the homepage. Microsoft just doesn’t get the value of a clean UX. DuckDuckGo results are pretty irrelevant the last time I tried them. There is nothing that comes close to their usability. To make the switchover, it has to be much much better than Go…

> Google simply has the best search product. The best available doesn't necessarily mean the best possible. And Google is far from it, and it's getting worse, not better.

I find the single hardest thing to search for these days is anything more than a few months old on YouTube... They hate older videos, it feels like. Beyond that, I keep seeing suggestions on new content from years ago... it's just weird.

I know it's not google proper, but I'd guess a significant number of their searches are specific to youtube.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#524

Ex-Google-Search engineer here, having also done some projects since leaving that involve data-mining publicly-available web documents. This proposal won't do very much. Indexing is the (relatively) easy part of building a search engine. CommonCrawl already indexes the top 3B+ pages on the web and makes it freely available on AWS. It costs about $50 to grep over it, $800 or so to run a moderately complex Hadoop job.…

If there were viable alternatives, people would shift over time. If I type in “ Pentagon” on Google, the first link is LinkedIn. DuckDuckGo doesn’t even list it at all. There’s countless examples where DuckDuckGo just can’t find basic information. DDG is just unreliable beyond it’s silly name.

What does a viable alternative look like?

I've been using Bing for the past few months; it's not great or terrible but is it "viable" enough for people to shift to over time? Or is it not viable because it's backed by a major corporation?

I'm sure there are search quirks with each engine but I've seen issues with Google too and yet it's the "devil we know" ... so people unconsciously work around them.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#525

Ex-Google-Search engineer here, having also done some projects since leaving that involve data-mining publicly-available web documents. This proposal won't do very much. Indexing is the (relatively) easy part of building a search engine. CommonCrawl already indexes the top 3B+ pages on the web and makes it freely available on AWS. It costs about $50 to grep over it, $800 or so to run a moderately complex Hadoop job.…

If there were viable alternatives, people would shift over time. If I type in “ Pentagon” on Google, the first link is LinkedIn. DuckDuckGo doesn’t even list it at all. There’s countless examples where DuckDuckGo just can’t find basic information. DDG is just unreliable beyond it’s silly name.

Unless you're searching in Russian, DDG is mostly a skin for Bing search results anyways. The major players in the search engine space are Google, Bing/MSN, Yandex, and Baidu - with the latter two being mostly language-specific.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#526

Earlier quoted context omitted.

> In order to get these protections, they have to behave a certain way. I'm not sure about this, but it's an argument I've heard. This is commonly repeated but not true. In fact, the opposite is true. https://www.law.cornell.edu/uscode/text/47/230 Scroll down to section C. 1. No provider or user of an interactive computer service shall be treated as the publisher or speaker of any information provided by another info…

I don't see how 2.a would stand in a courtroom. What is 'good faith' and 'otherwise objectionable' in this context? Is removing all content from a religious group afforded in this definition, as long as the provider considers it objectionable? 'whether or not such material is constitutionally protected' has almost no meaning because almost all speech is protected. Another interesting point is, while it's allowable by…

I don't think it makes sense for us to debate "what would stand in a courtroom" since this has been the law of the land for more than two decades and has been able to "stand" just fine so far. We could debate whether or not the law should be repealed or amended, but the consequences of the law as is are pretty clear: online platforms are not responsible for content posted by a 3rd party and are free to curate content on their platform at their own discretion.

> is it allowable to promote content you prefer over other content? Or de-promote the content but not remove it?

Legally it is allowable since the law makes no distinction between "promotion" and "curation".

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#527

Earlier quoted context omitted.

> Google simply has the best search product. The best available doesn't necessarily mean the best possible. And Google is far from it, and it's getting worse, not better.

I've definitely noticed a decline in quality from Google results over the past few years in particular. I don't know if that's because SEO has gotten control of the results of if Google's algo is shoving lower quality up higher for revenue but it's become difficult. Using a bit of Google-fu I'm usually able to find what I need quickly but it's still more of a hassle than it used to be.

There's exponentially more background noise than there used to be

It's easier to return the most relevant 10 results when there's only 10 thousand options than when there's 10 trillion options with 10 thousand new ones created every day.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#528

Earlier quoted context omitted.

The data was about 55TB of compressed HTML last I looked, so that's about 70 r5a.24xlarge instances, each going for $5.424/hour, so about $350/hour or $250K/month. That's not cheap, and definitely not something you'd put on your personal credit card, but it's well within the range of a seed-funded startup. Sizes may vary a bit depending upon the exact index format, but that should be a rough ballpark. With batch jobs…

"Remember, search touches (nearly) every indexed document on every query" - wait, why does that happen? Doesn't it only touch ones with at least one of the search terms in, or stemmed/varied words relating to some of the terms? And does that via an index?

I struggled with how to word that in a way that's both true, understandable, and doesn't give away any proprietary information. Added "indexed" to clarify but I didn't fix up the numbers, so they're likely an overestimate.

Basically, yes, it uses an index and touches only documents that appear in one of the relevant posting lists. However, after stemming, spell-correcting, synonyms, and a number of other expansions I'm not at liberty to discuss, there can be a lot of query terms that it needs to look through, covering a significant portion of the index. Each one of these needs to be scored (well, sorta - there are various tricks you can use to avoid scoring some docs, which again I'm not at liberty to discuss), and it's usually beneficial to merge the scores only after they have been computed for all query terms, because you have more information about context available then.

There's a reason Google uses an in-memory index: it gives you a lot more flexibility about what information you can use to score documents at query time, which in turn lets you use more of the query as context. With an on-disk index you basically have to precompute scores for each term and can only merge them with simple arithmetic formulas.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#529
post #361

Ex-Google-Search engineer here, having also done some projects since leaving that involve data-mining publicly-available web documents. This proposal won't do very much. Indexing is the (relatively) easy part of building a search engine. CommonCrawl already indexes the top 3B+ pages on the web and makes it freely available on AWS. It costs about $50 to grep over it, $800 or so to run a moderately complex Hadoop job.…

This ^ times a 1000. Google simply has the best search product. They invest in it like crazy. I’ve tried bing multiple times. It’s slow, it spams msn ads in your face on the homepage. Microsoft just doesn’t get the value of a clean UX. DuckDuckGo results are pretty irrelevant the last time I tried them. There is nothing that comes close to their usability. To make the switchover, it has to be much much better than Go…

Microsoft thinks what they have is Clean UX.....

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#530

Ex-Google-Search engineer here, having also done some projects since leaving that involve data-mining publicly-available web documents. This proposal won't do very much. Indexing is the (relatively) easy part of building a search engine. CommonCrawl already indexes the top 3B+ pages on the web and makes it freely available on AWS. It costs about $50 to grep over it, $800 or so to run a moderately complex Hadoop job.…

>The comments here that PageRank is Google's secret sauce also aren't really true - Google hasn't used PageRank since 2006.

That's quite a claim considering they were reporting PageRank in their toolbar until 2016, and toolbar PageRank was visible in Google Directory until 2011.

Are you talking about PageRank from the original patent?

Post reply on HN