Live data from Hacker News

To Break Google’s Monopoly on Search, Make Its Index Public

bloomberg.com

501–510 of 630 posts

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#502

Earlier quoted context omitted.

> Google simply has the best search product. The best available doesn't necessarily mean the best possible. And Google is far from it, and it's getting worse, not better.

I've definitely noticed a decline in quality from Google results over the past few years in particular. I don't know if that's because SEO has gotten control of the results of if Google's algo is shoving lower quality up higher for revenue but it's become difficult. Using a bit of Google-fu I'm usually able to find what I need quickly but it's still more of a hassle than it used to be.

I don't know why you're getting downvoted, because the quality has 100% tanked over the last few years. I agree that there may be some selection bias between us, but it's at least got some of my normie non-technical friends commenting about it, so it's not completely without merit. I have a couple of theories, one of them is also a warning.

First, I think search results at Google have gotten worse because people are not actually good at finding the best example of what they're looking for. People go with whatever query result exceeds some minimum threshold. This means when Google looks at what people "land on" (e.g. something like the last link of 5 they clicked from the search page, and then which they spend the most time on according to that page's Google Analytics or whatever), they aren't optimizing for what's best, they're optimizing for what is the minimum acceptable result. And so what's happening is years and years of cumulative "Well, I suppose that's good enough" culminating in a perceptible drop in search result quality overall.

Second, Google has clearly been giving greater weight to results that are more recent. You'd think this would improve the quality of the results which "survive the test of time" but again, Google isn't optimizing for "best" results, they're optimizing for "the result which sucks the least among the top 3-5 actual non-ad results people might manage to look at before they are satisfied". So this has the effect of crowding out older results which are actually better, but which don't get shown as much because newer results have temporal weight.

My warning is this, too, which you've surely noticed: Google search has created a "consciousness" of the internet, and in the 90s it used to be that digitizing something was kind of like "it'll be here forever" and for some reason people still today think putting something online gives it some kind of temporal longevity, which it absolutely does not have. I did a big research project at the end of the last decade, and I was looking for links specifically from the turn of the century. And even in 2009, they were incredibly hard to find, and suffered immensely from bitrot, with links not working, and leaning heavily on archive.org. Google has been and is amplifying this tremendously, by twiddling the knob to give more recent results a positive weight in search. Google makes a shitload of money from mass media content companies (e.g. Buzzfeed) and whatever other sources meet the minimum-acceptable-threshold for some query, versus linking to some old university or personal blog site which has no ads whatsoever. So the span of accessible knowledge has greatly shrunk over the last few years. Not only has the playing field of mass media and social media companies shrunk, but the older stuff isn't even accessible anymore. So we're being forced once more into a "television" kind of attention span, by Google, because of ads.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#503

This seems a bit silly - it's not like Google's search results are that much better than Bing or DuckDuckGo (or better at all most of the time). Google has Google Colab, Chrome, and Android, and the whole Docs ecosystem, and they've integrated those things together pretty well while still being loosely coupled enough to switch out any part of the ecosystem with a competitor's product. Is there a reason to break up Go…

Colab? I don't think that google's position is held in power by that.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#504
post #474

Earlier quoted context omitted.

The first sentence of the linked New York Times story: > Google on Tuesday acknowledged to state officials that it had violated people’s privacy during its Street View mapping project when it casually scooped up passwords, e-mail and other personal information from unsuspecting computer users. That answers your first three paragraphs. There’s no “if” to their lying and privacy invasions. They’ve been caught and admit…

What? You seem to be misunderstanding my statements. My first points were about the streetview product. Scooping up passwords is obviously not the intent of that product, maybe that was an error or they changed the core product at some point? I can't read the paywalled article. I'm not suggesting non-technical users create products... you're reading so far out of context. Just because user X can't create a new produc…

> Just because user X can't create a new product does not mean that we should place sanctions on company Y. (…) in this post you just illogically connect a bunch of dots.

That is an insane extrapolation, and the reason I don’t want to continue the conversation with you: you’re answering points I’m not making. I haven’t even hinted at sanctions; I have no idea where you’re getting that from.

> But a fact is still no one is forcing you to use these products

And I don’t use them. I hoped that by continuing to mention non-technical users you’d get it, but this was never about me. You keep bringing up that argument, but read what you replied to in the first post — I recounted the experience of non-technical people I know, not my experience. Stop telling me I have a choice; the point is not us, it’s non-technical users who don’t have the knowledge to make informed choices!

> That sounds like a rationality of a completely one-sided biased individual in itself, respectfully.

Believe what you want. I just don’t want to keep wasting my night arguing with someone that started a discussion but refuses to address the points originally made. Why reply, then?

Maybe I’m not explaining myself well enough, or in the correct way for you to understand, or maybe you’re the one not grasping what I mean. It doesn’t really matter where the problem lies, just that it’s clearly not working.

Maybe if we ever meet in person we can resume this conversation, but tonight it’s not being productive, so I genuinely wish you a good week and sign out here.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#505

Earlier quoted context omitted.

> Indexing the top billion pages or so won't take as long as people think. This is what makes me wonder why we don't have a LOT of competing search engines. Perhaps i'm vastly under-estimating the technology and difficulty (I could well be - it's not my domain) but it surely it can't be THAT hard to spawn Google-like weighted crawl-based search results? It's a long-since solved problem - heck, pageRank's first iterat…

Most likely answer: lack of diversity in revenue models. Outside of ad revenue, search has always been seen as something of a "charity" effort for the internet. It's "boring" infrastructure work that can be critically useful but doesn't really make money directly on its own. No one wants to pay a "search toll" and there's no government agency in the world that the internet would trust as a neutral index to run it as…

Which begs the question, if adblock makes advertising based models go the way of the dodo, what happens to search?

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#506
Why does everyone assume Google is automated? A single website containing a small amount of knowledge or even a leaked archive of Diebold emails will show lower on Google search results than a web forum or some listicle.

The real web crawlers are people who surf the web, Google likely takes information from analytics and outbound link tracking to determine the most popular (not the most informational) website.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#507
post #426

Earlier quoted context omitted.

> Also people forget that the creepy stuff Google does is super useful . For the same reasons you’re exalting them, I have non-technical friends who asked me how Google knows so much about them (and suggestions on how to avoid it) because they found it too creepy. I don’t think people forget Google’s results are useful; some just think they’re more creepy than valuable. You seem to have picked your side in that (im)b…

I wish interfaces were more straight up about their intentions and made it easier to implement account level partitions. For work I love Google's magic tracking effects, but at 1 am, hell no.

You can have multiple identities in chrome[0], even guest identities.

[0] https://support.google.com/chrome/answer/2364824

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#508

Earlier quoted context omitted.

I think specifically in the case of platform vs publisher, there is federal law that gives some sort of extra legal protection to platforms. In order to get these protections, they have to behave a certain way. I'm not sure about this, but it's an argument I've heard. > Liable is incredibly hard to prove in court, and usually cases are taken in civil court not criminal court. Going to court is expensive, even if you…

> In order to get these protections, they have to behave a certain way. I'm not sure about this, but it's an argument I've heard. This is commonly repeated but not true. In fact, the opposite is true. https://www.law.cornell.edu/uscode/text/47/230 Scroll down to section C. 1. No provider or user of an interactive computer service shall be treated as the publisher or speaker of any information provided by another info…

I don't see how 2.a would stand in a courtroom.

What is 'good faith' and 'otherwise objectionable' in this context? Is removing all content from a religious group afforded in this definition, as long as the provider considers it objectionable?

'whether or not such material is constitutionally protected' has almost no meaning because almost all speech is protected.

Another interesting point is, while it's allowable by the law to restrict content you deem objectionable, is it allowable to promote content you prefer over other content? Or de-promote the content but not remove it?

Is it acceptable to remove religious content with no explanation as to why? Is it acceptable to ban or demonetize a religious group without explanation?

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#509
post #426

Earlier quoted context omitted.

> 1) a record of searches and user clicks for the past 20 years From what I can tell, Google cares a lot more about recency. When I switch over to a new framework or language, search results are pretty bad for the first week, horrible actually as Google thinks I am still using /other language/. I have to keep appending the language / framework name to my queries. After a week or so? The results are pure magic. I can…

> Also people forget that the creepy stuff Google does is super useful . For the same reasons you’re exalting them, I have non-technical friends who asked me how Google knows so much about them (and suggestions on how to avoid it) because they found it too creepy. I don’t think people forget Google’s results are useful; some just think they’re more creepy than valuable. You seem to have picked your side in that (im)b…

I don’t think people forget Google’s results are useful; some just think they’re more creepy than valuable. You seem to have picked your side in that (im)balance, and other people prefer the other side.

Just as a general observation without taking either side:

People routinely fail to recognize both sides of a particular thing. It's why we have sayings like "You don't know what you've got til it's gone."

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#510

Ex-Google-Search engineer here, having also done some projects since leaving that involve data-mining publicly-available web documents. This proposal won't do very much. Indexing is the (relatively) easy part of building a search engine. CommonCrawl already indexes the top 3B+ pages on the web and makes it freely available on AWS. It costs about $50 to grep over it, $800 or so to run a moderately complex Hadoop job.…

Actually the omnibox made it really easy to switch to ddg. With an occasional fallback to google. I have no problem with advertising etc. but the tracking and selling of data is such an idiotic thing. We as consumers should have a global internet-law, and be reimbursed for data leaks or usage outside the scope of the application. By no problem with ads I mean the original ads of google. Very clear they were ads and n…

I think that the "fallback to Google" might actually tend to diminish consumer confidence in DDG. Every time you use it, you basically say to yourself "$newRiskyStrategy fails sometimes, we still need $oldReliableStrategy".

Instead, what might help DDG is a plugin that detects when you go past the first or second page of Google search results, and suggests that you might get better results on DDG. It's a little intrusive, but the mental nudge becomes "$oldReliableStrategy has flaws, try $newRiskyStrategy". You get a positive emotional interaction with DDG rather than "forcing" yourself to use it all of the time and "failing back" to Google.

Post reply on HN