Live data from Hacker News

To Break Google’s Monopoly on Search, Make Its Index Public

bloomberg.com

311–320 of 630 posts

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#311
post #233

Earlier quoted context omitted.

> 1) a record of searches and user clicks for the past 20 years If a government was serious about getting more players in the search industry, they would force Google (and all other players) to make this data public. Simply say "All user-behaviour data used to improve the service must be freely published". Make the law apply to any web service with more than 20 million users globally so small businesses aren't burden…

> If the data cannot be published for privacy reasons, the private parts must be seperated and not used by google or it's competitors. As a user that notices the impact of this data: please no, thanks though. Have you ever visited youtube's home page in incognito mode? It's... bad. Really bad. Not allowing any company to use this (obviously very private) information in ranking would simply make their products suck, h…

>Have you ever visited youtube's home page in incognito mode?

Do you like the personalized recommendations because of channel subscriptions?

I always get the "anonymous default" home page with YouTube and don't care. The home page is just a wasted load before I can start typing in the search bar. As a bonus, staying incognito means all the videos on the right-side panel are related to the current video. Not related to a music video I have playing in another tab.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#312
post #287
post #81

Earlier quoted context omitted.

> Yes, Google has a huge index, but most queries aren't in the long tail. I'm not quite sure about that. 15% of Google searches per day are unique, as in, Google has never seen them before. [1]. That's quite an insane number. [1] https://searchengineland.com/google-reaffirms-15-searches-ne...

> 15% of Google searches per day are unique, as in, Google has never seen them before. That is impossible, and therefore wrong (I'm wrong, please see below). To know if a search is unique, as in Google has never seen them before, Google must be able to decide if a query it receives was seen before or not. Even if we assume Google needed only one bit for each message it has ever seen, and assuming it only saw 15% of n…

I don't think it's necessarily impossible to calculate. Using probabilistic data structures arranged in a clever way, it's likely possible to calculate with some degree of accuracy.

I haven't thought this through, but take all the queries as they're made and create a bloom filter for every hour of searches. Depending when this process was started, an analytics group could then take a day of unique searches, and run them against this probabilistic history, and get a reasonable estimation with low error. Although the people who work on this sort of thing probably know it far better than I.

The real question though might be assuming the 15% is right, do we care about those 15%, are they typo's that don't merge, are they semantically different, are they bots search for dates or hashes, etc.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#313
post #39

Earlier quoted context omitted.

What about the computing and infrastructure resources that Google dedicated to build the index, and continues to dedicate to keep the index up to date? This is a failed model. There is no incentive for Google to continue to update the public version, and it would quickly fall out of date while Google focuses on their own internal copy. It’s a solution suggested by lawmakers who fundamentally don’t understand how comp…

> What about the computing and infrastructure resources that Google dedicated to build the index How is this different from the 1956 consent decree? AT&T spent money developing its patents. But on the basis of longstanding law around public interest, it was forced to license them to third parties. (Note: not give them away.) > and continues to dedicate to keep the index up to date? The author explicitly contemplates,…

Conflating AT&T to Google is incredibly misleading. AT&T had become a defacto monopoly approx. 20 years prior to that consent decree, not to mention being subjected to an anti-trust lawsuit ten years prior to the decree.

The article makes the conceit of equating an online data store to the telephone infrastructure. Google is not preventing other companies from indexing the internet in any way, shape or form. There is no barrier to another company building a similar data store from scratch using the existing internet infrastructure. On the other hand, laying out your own wires has physical limitations, especially when another company has claimed the most optimal route. In such a scenario, your costs will always exceed theirs (more infrastructure to reach the same consumers).

Regarding forcing Google to use the same index: How does DoJ plan to enforce this? How would DoJ ever be sure that Google is in compliance? I once again assert that the law magically assumes new technology to solve old technology problems without fundamentally understanding how computers work. Computers are copy-on-write by design. Delete always requires an extra step, and the verification that only one copy exists is impossible to make in the digital domain.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#314

Caveat: The author is not a technologists Robert Epstein (born June 19, 1953) is an American psychologist, professor, author, and journalist. He earned his Ph.D. in psychology at Harvard University in 1981, was editor in chief of Psychology Today, He has also made some questionable claims about google manipulating search results to favor Hillary Clinton. https://en.wikipedia.org/wiki/Robert_Epstein#cite_note-15 His r…

Just FYI the completion results in the omnibox have little to do with the search engine results. Clearly the search engine produces millions of hits for “Hillary Clinton emails”. The completions are a completely separate system based on what people type in the box, not what’s in the index, and it’s laser-focused on producing interactive results.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#315
post #72

Earlier quoted context omitted.

Foreign propaganda usually doesn't label itself as such. Nor does this article have to be from such a source. Merely inspired by the noise being made

Are you suggesting that we should discard criticism of google because Russia and China is anti google? China does not like Trump either, so should we not criticize trump because of China? And I don't even get why you say China and Russia is anti Google. They are anti-information more than anti-google. If google allows them to control what information people get they will have no problem with Google. They are further…

I'm not suggesting that. Nor am I saying China & Russia are anti Google. I was only responding to why your statement "article doesn't mention X,Y,Z" was missing the mark on taf2's argument

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#316
post #152

Earlier quoted context omitted.

I would assess Google (& FB's) "crown jewel" as, ultimately, their market share, which is related to your points... and causation runs both ways. The user data helps/ed Google create the superior UX, as you say. The reach is what makes Google & FB valuable to advertisers. A search engine with 0.1% of Google's user volume cannot charge advertisers 0.1% of Google's as revenue. Returns to scale/reach/market-share are ve…

'Break Google up' would mean you'd have: * an Office suite / enterprise company (Google Cloud + Docs + Gmail + Business) * a phone company (Android) * a search company (Google Search + Advertisement) * and a media company (Google Play Movies, Music, Books and YouTube) The names would probably become different in time, but you get the gist. Amazon and Microsoft could be broken up much the same way, in neat categorical…

Breaking up companies like Google, Amazon, and Microsoft are just not gonna happen in 2019 where huge, global mega-corps are the only way to compete outside of small local markets.

Even though a lot of these corporations build offices, hire non-Americans, and pay tons of foreign taxes in countries in which they do business, the main executives and talent still live in the US, the IP is developed here, and the majority of profits end up back in the home country.

It's better for everyone who actually matters - shareholders, intel agencies, government officials, associated businesses, etc - that these companies remain large and globally dominant, even if it screws over US citizens by having to pay the monopoly taxes and suffer the privacy invasions. We're an insignificant sacrifice in the decision-makers' minds.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#317
post #193

Earlier quoted context omitted.

Some people are arguing that catering search results or what content is allowed on a platform to a specific set of political views makes those platforms publishers rather than mere platforms. Apparently this also has some implications in some political campaign laws I don't really understand. I think we're heading towards political ideologies having the same protection as religious institutions (which IMO, are exactl…

I agree that political ideologies are similar to religions and that protections may be extended to them (although they should already be covered by what is in the constitution). What I don't follow is why it matters that platforms are becoming publishers. The government has never had the right to meddle in what publishers decide to publish (with exceptions for regulations on pornography and classified information). I…

I think it matters in the context of liability.

For instance, if you're a newspaper, and you publish "Politician X is a rapist" and you don't have a way to back up that allegation in court you're liable. As a platform, you're not liable for the content your users post.

If you are allowing people to say "Politician X is a rapist" but not allowing people to say "Politician Y is a rapist", you've crossed into the land of publisher in some people's minds.

Simply being a platform on the web is 'doing business' if you're getting revenue from that operation. Can a platform specifically reject a religious group from their site? Can a platform demonetize a religious group?

I personally don't believe in protected groups, but if we're going to have protected groups, we should apply the same protection to all groups, equally, even if the group size is 1.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#319
post #61

Okay, this is a relatively serious proposal to require Google to allow API access to its search index, with the premise that it would democratize the search engine ecosystem. There are some issues with the regulations he proposes (you have to allow throttling to prevent DDoS attacks, and you can't let anyone with API access add content to prevent garbage results), but it's roughly feasible. The main problem is, I thi…

>2) 20 years of experience fighting SEO spam. Tangential - but does anyone else feel that google results are useless a lot of the time? If you search for something, you will get 100% SEO optimized shitty ad-ridden blog/commercial pages giving surface level info about what you searched about. I find for programming/IT topics its pretty good, but for other topics it is horrible. Unless you are very specific with your s…

I agree with this. Most searches give me almost a whole page of ads and stuff up top before the things I’m interested in start showing up way down at the bottom of the page, and even then the results are often spam.

I’ve been using DuckDuckGo and have found I have this problem less. I don’t always find what I mean on DDG, as of now I’d say Google is still better if you’re not sure exactly what you’re looking for is called, but if you know the keywords you need DDG is often better.

Re: To Break Google’s Monopoly on Search, Make Its Index Public

#320
post #61

Okay, this is a relatively serious proposal to require Google to allow API access to its search index, with the premise that it would democratize the search engine ecosystem. There are some issues with the regulations he proposes (you have to allow throttling to prevent DDoS attacks, and you can't let anyone with API access add content to prevent garbage results), but it's roughly feasible. The main problem is, I thi…

> Indexing the top billion pages or so won't take as long as people think. This is what makes me wonder why we don't have a LOT of competing search engines. Perhaps i'm vastly under-estimating the technology and difficulty (I could well be - it's not my domain) but it surely it can't be THAT hard to spawn Google-like weighted crawl-based search results? It's a long-since solved problem - heck, pageRank's first iterat…

It's not the 'raw' search itself. It's the billions (trillions) of queries they've captured: Person X searches for query Y and clicks on result Z.

This is far more valuable than the general page rank algorithms that were initially developed and have already been duplicated many times in academia and business.

Post reply on HN