Live data from Hacker News

To Break Google’s Monopoly on Search, Make Its Index Public

bloomberg.com

251–260 of 630 posts

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#251
post #100
post #92

Earlier quoted context omitted.

Also - just the amount of people who would game the system afterwards all the search results would be utterly useless.

Ye and anyway you could return x10 worse results than Google's current results and still become the new dominant search engine if you've got infinite $B a year to outbid Google to be the default on browsers and operating systems and a big salesforce to onboard the advertisers. Man this is the exact playbook they used to become the search engine in the first place. They literally were Yahoo's search bar at some point…

Microsoft has infinite money, pays users with rewards directly, and has barely gotten any marketshare with Bing.

In order to compete, you would need good results and good performance for your organic results and your ad program in addition to a huge budget for traffic aquisition and patience and ability to execute on strategy consistently over multiple years.

Also, keep in mind that people will judge based on perceived quality, not objective quality. Simply being shown as a Google result increases the perceived quality for most results -- in order to be seen as equal, the competitor will need to have consistently better results.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#252

Earlier quoted context omitted.

How does your performance scale? If it scales linearly in records, you're going to have trouble if you're only below 200ms on 1k records. If it scales logarithmically, there are still a couple orders of magnitude between a comprehensive index like Google's and your 1k record index, which is still going to give you trouble. In short, 200ms is slow, not fast. EDIT: Just to clarify, I'm not saying it's necessarily easy…

Thanks for the insight; indexing the entire web is not a goal of dindex. Instead, it's designed to work like DNS where you can host your own server + have it federate queries to other servers if it doesn't have records which match a query. Basically it moves control to the client; if the client wants to query servers across the globe there's nothing stopping them, but by default you only query servers in your config…

But yea while performance is a huge goal of mine, at the moment I don't even cache compiled regular expressions. It's definitely an alpha-stage project.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#253
post #206

From the article: "But what about those nasty filter bubbles that trap people in narrow worlds of information? Making Google’s index public doesn’t solve that problem, but it shrinks it to nonthreatening proportions. At the moment, it’s entirely up to Google to determine which bubble you’re in, which search suggestions you receive, and which search results appear at the top of the list; that’s the stuff of worldwide…

The other problem with this is that it still can't change human nature. Ok, so this plan is implemented and any site can serve google results and order them as they want with an API. People are still going to go to their favorite far right or far left outlets, which can now access google results and show only the articles that they know their users want to see. The "filter bubble" problem could even be worse than it…

That was my first thought. So this is a thing, and now we have dedicated search engines to showing you the 'true' search results from either the left or right. And it's that much easier to stay completely enveloped in the echo chamber.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#254
post #81
post #61

Okay, this is a relatively serious proposal to require Google to allow API access to its search index, with the premise that it would democratize the search engine ecosystem. There are some issues with the regulations he proposes (you have to allow throttling to prevent DDoS attacks, and you can't let anyone with API access add content to prevent garbage results), but it's roughly feasible. The main problem is, I thi…

> Yes, Google has a huge index, but most queries aren't in the long tail. I'm not quite sure about that. 15% of Google searches per day are unique, as in, Google has never seen them before. [1]. That's quite an insane number. [1] https://searchengineland.com/google-reaffirms-15-searches-ne...

Could this be explained by supposing that people are just searching for current events, sometimes national, sometimes international, sometimes very local? If so, you really wouldn't need much indexed to handle those queries. I imagine many queries are also just overly verbose and sentence-length, which artificially inflates the number of unique queries which are actually seeking roughly the same pages.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#255
post #170

Earlier quoted context omitted.

Google's claim that the algorithm is generic is demonstrably false. Type in "hillary clinton e" and there is no suggestion for "email", type "donald trump e" and email is the first suggestion. Given the news content that we know is out there, that can only be the result of adjusting the results for clinton specifically (if anything, we would not expect "email" to be autocompleted for trump). This is not research that…

This is not "research" period. Using one arbitrary search comparison to draw conclusions about the nature of a system that processes billions of queries a day is pretty weak. Additionally, I don't get the same results you do. "hillary clinton e" does not bring up emails, nor does "donald trump e" bring up emails (the first results I see are election, education, england visit, ex wife). I'm not ruling out the possibil…

That's the point: this is not research, but whatever is going on at Google, the explanation has to account for examples like these. It's simply one observation that you cannot discount.

I just tried searching again a few times with new private windows, and "email" alternates between first and fourth suggestion for trump. But the more important point is the absence of the suggestion for clinton: we know it's been in the news extensively, we know people searched for this phrase a lot, and now "email" has been removed from the suggestions only for Clinton. I tried searching a few more U.S. politicians, and for all of them "e" autosuggests "email" somewhere between first and fourth place. So the complete absence for Clinton does not look like a generic algorithm change.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#256

Earlier quoted context omitted.

That which the market does not naturally provide... does one just morally expect?

Two points. The slap on the wrist from the DOJ didn’t lead to decreasing dominance on Microsoft. Google/Facebook/Apple competing did. A search engine is not life or death. Google doesn’t stop anyone from creating a website and reaching consumers other ways.

>A search engine is not life or death. Google doesn’t stop anyone from creating a website and reaching consumers other ways.

As someone outside of the tech bubble, this statement always confuses me. I see it once a week or so on this site. Someone will say 'well, company x is doing y, what's to stop you from just disrupting them and working around?"

Because people get their information in consistent, and predictable ways. And one of those main ways is to search for it using a popular website, so popular it's literally called 'googling' something.

If you are an unknown, a search engine is literally life and death for your company, from what I can tell.

Again, I'm outside the tech sphere, so maybe I'm just ignorant. But I don't think so.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#257
post #81

Earlier quoted context omitted.

> Yes, Google has a huge index, but most queries aren't in the long tail. I'm not quite sure about that. 15% of Google searches per day are unique, as in, Google has never seen them before. [1]. That's quite an insane number. [1] https://searchengineland.com/google-reaffirms-15-searches-ne...

Now I feel bad for putting gibberish like jsjsjdkktkwoapaoalf in my address bar and searching Google to test if my internet is working..

I do that all the time, I wonder how common that is?

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#258
post #203

Earlier quoted context omitted.

So... I think there are two issues with this. (1) This doesn't actually reduce market share, since each of these are basically different market categories. (2) Almost all the revenue is from search. That company is the revenue generating arm for the other ones.

Yes, my thought was that by breaking everything off from everything else, these silo'd services would suddenly have to compete with the rest of their market at fair terms, instead of being propped up massively by other division(s), and thus would lose marketshare to a multitude of fresh and established competitors. You are right though, it doesn't deal with the dominance of the search directly. My hope is a complimen…

>You are right though, it doesn't deal with the dominance of the search directly. My hope is a complimentary effect to the above also happens: Google no longer gets gobs of personal data from its other services, allowing other search engines to approach its efficacy.

I'm still not sure how this would work on Apple though, since their main differentiator is their design sensibilities and integration rather than their platform monopolies.

I guess iMessage and the App Store do rely on monopoly rents, but I can't think of any way to sever those links without making the iOS platform less secure.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#259

No sure if relevant but I've hated google's search for a unique reason: I find it horribly slow, especially compared to what it used to be. It's so slow, I've designed + partially implemented an alternative for my own use: https://github.com/Jeffrey-P-McAteer/dindex I've only tested with 1000 records, but the query times are all <200ms.

Maybe it's your part of the world, but my TTFB to google search pages is less than 150ms and the "generated in x seconds" is usually less than 1 second. That's pretty good for an index searching effectively every public internet page.

I measure from when I hit enter to when my screen is full of results, and just _rendering_ google.com takes a full second on my macbook with 8gb ram and an i5 processor. It's so bad I have a shell script which forces my processor into the 3200mhz range when I'm on my browser workspace. When I'm not looking at my browser (+no downloading files, no audio) the same script sends a SIGSTOP to it so it isn't eating CPU cycles while I'm writing code.
Post reply on HN