Live data from Hacker News

Data is at the heart of search, but who has access to it?

andreasgal.com

61–70 of 88 posts

Re: Data is at the heart of search, but who has access to it?

#61
I think the author's perspective is skewed in order to stay in line with the title. Here's an example of why I say that:

The author states that "For some 90% of searches, a modern search engine analyzes and learns from past queries, rather than searching the Web itself, to deliver the most relevant results." This may be true in some types of searches but overall, I think the statement is misleading.

Rather, it's better to think of it like this: One important part of the algorithmic process involves constantly crawling the web and updating the index with new information. (Important / frequently-updated web sites may get crawled all day every day, while ones that are less important may get crawled only weekly or monthly). Meanwhile, another part of the algorithmic process constantly analyzes new info discovered in the crawl and combines it with, as the author-mentioned, click-through data learned from past queries.

The answers to many queries don't change, while the answers to many other queries deserve freshness. For example, I'm quite certain Einstein's date of birth hasn't changed in quite a while, but his theory of relativity is in constant discussion and there is always new information and new queries pertaining to it. As a result, there is not much need for a search engine to go digging for the latest info on an "einstein's birthday" query, but it's to everyone's advantage that Google is able to identify which pages on the web deserve priority crawling and that Google has retrieved and incorporated the fresh info those pages contain into its index when it comes to a topical type of query like "diffraction of light with quantum physics".

In the end, the results to every query depend on info gathered from the web and user data helps refine the results. Info that is more static can be prioritized with more input from click-through data, while new information found on the web must rely more on Google's artificial intelligence to push it up in front of searchers.

Another reason that that "90%" statement sticks out to me is that there is a fairly often-used factoid tossed around industry experts that between "6% to 20% of queries that get asked every day have never been asked before." Google can't rely heavily on past query data for all of these type of searches.

Re: Data is at the heart of search, but who has access to it?

#62

Sigh, this is incorrect. edit: incorrect is perhaps too strong, it is incomplete. While it is true that click tracking can be used as a relevance signal, the people who were really pissed off when the data stream got dumped were advertisers who wanted to buy AdWords. That was a very simple system, pay someone for clickstream data, extract trending queries, front those with AdWord buys to get your page on the top of G…

But where's the competitive ecosystem in search? Innovation in search is restricted to few hundred people in Mountain View. And that's a tragedy.

What Google did for innovation in smartphone\tablet\browser they have gone and done the opposite for search.

Re: Data is at the heart of search, but who has access to it?

#63

I think the author's perspective is skewed in order to stay in line with the title. Here's an example of why I say that: The author states that "For some 90% of searches, a modern search engine analyzes and learns from past queries, rather than searching the Web itself, to deliver the most relevant results." This may be true in some types of searches but overall, I think the statement is misleading. Rather, it's bett…

You're vastly underestimating the uniqueness of search queries these days. Various sources within Google have said that 25% to 50% of queries entered into Google have never been seen before at all.

Re: Data is at the heart of search, but who has access to it?

#64
post #10

So the whole push for SSL/https from Google has been opportunistic rather than good practice. I mean why would a search engine go as far as to make SSL a ranking signal?

It could be opportunistic and a good practice. Users do benefit from sites that offer SSL. It's just that Google benefits too.

Re: Data is at the heart of search, but who has access to it?

#65
post #10

So the whole push for SSL/https from Google has been opportunistic rather than good practice. I mean why would a search engine go as far as to make SSL a ranking signal?

> I mean why would a search engine go as far as to make SSL a ranking signal?

Because any number of 3rd parties have been injecting their ads and other crap as MITMs. SSL is a better, but not foolproof way to make sure the content you get was the content served by the remote server.

Re: Data is at the heart of search, but who has access to it?

#66

Sigh, this is incorrect. edit: incorrect is perhaps too strong, it is incomplete. While it is true that click tracking can be used as a relevance signal, the people who were really pissed off when the data stream got dumped were advertisers who wanted to buy AdWords. That was a very simple system, pay someone for clickstream data, extract trending queries, front those with AdWord buys to get your page on the top of G…

Chuck, while blekko is a great search engine(especially due to custom search), it's clear that it is very different quality wise from Google.Same for Bing - it's not upto Google.And not for the lack of trying or money(bing).

So how do you think Google is succeeding so well, if it's not click stream data? and why can't it be maybe a combination of things that strongly depends on click stream data that others couldn't copy?

Re: Data is at the heart of search, but who has access to it?

#67
post #48

Earlier quoted context omitted.

I get this completely, although let's say Microsoft just ignored their requests for compliance. Would the EU seriously dare to ban Windows? I feel like they'd get outrage from their own locals and topple their own economy if they banned Windows, so they probably wouldn't. Therefore, does Microsoft need to care? Could they just sit around in Redmond and keep developing as long as the US doesn't care? As for websites,…

They would fine Microsoft. If Microsoft didn't pay they would have their assets seized. Or their credit rating damaged. Their credit rating drop could then put them as junk status and thus mutual funds would have to divest from Microsoft. Microsoft share price would be negatively affected. NB: I am not a lawyer or economist.

[deleted]

Re: Data is at the heart of search, but who has access to it?

#68
post #48

Earlier quoted context omitted.

I get this completely, although let's say Microsoft just ignored their requests for compliance. Would the EU seriously dare to ban Windows? I feel like they'd get outrage from their own locals and topple their own economy if they banned Windows, so they probably wouldn't. Therefore, does Microsoft need to care? Could they just sit around in Redmond and keep developing as long as the US doesn't care? As for websites,…

They would fine Microsoft. If Microsoft didn't pay they would have their assets seized. Or their credit rating damaged. Their credit rating drop could then put them as junk status and thus mutual funds would have to divest from Microsoft. Microsoft share price would be negatively affected. NB: I am not a lawyer or economist.

[deleted]

Re: Data is at the heart of search, but who has access to it?

#69
post #66

Sigh, this is incorrect. edit: incorrect is perhaps too strong, it is incomplete. While it is true that click tracking can be used as a relevance signal, the people who were really pissed off when the data stream got dumped were advertisers who wanted to buy AdWords. That was a very simple system, pay someone for clickstream data, extract trending queries, front those with AdWord buys to get your page on the top of G…

Chuck, while blekko is a great search engine(especially due to custom search), it's clear that it is very different quality wise from Google.Same for Bing - it's not upto Google.And not for the lack of trying or money(bing). So how do you think Google is succeeding so well, if it's not click stream data? and why can't it be maybe a combination of things that strongly depends on click stream data that others couldn't…

Actually if you do double blind tests you will find that Bing and Google are indistinguishable. We did this at Blekko earlier with our "3 card monte" gambit where you did a query, got back blekko, bing and google results, and got to pick the one with the "best" results for your query. Blekko usually won if it was query we had a slashtag for or if it was a "highly contested" query (lots of ad spend like "no fee credit card" or "cheapest insurance") In the former case our curation meant that more results were appropriate, and in the latter case our spam filtration left us with better results. If it was a general query for which we didn't have a category for, and it wasn't highly contested, google and bing split the results, often 40/40/20 sometimes as low as 35/35/30. And if it was a long tail query like "turnip growing in south philidelphia" or something very specific with few sites associtated with it, and we didn't have it in a slashtag, Google would "win" those. Microsoft borrowed our idea and did their whole "bing and decide" campaign.

Many people realize that if you put Google ads on Bing's results and Bing's ads on Google results the profitability would switch (not that I am entirely sure what that says other than having a credible search engine and top end Ad inventory is required to make excess money in search)

It will be interesting to see if Marissa gets back into the game with Yahoo when their agreement to use Bing results for Yahoo searches expires.

The interesting linkage is that you can't sell search advertising unless people send the search request to you, and if you're not the most common place that people search, you're unlikely to get first shot at advertising. You can "buy" traffic (that is called Paid Distribution) by putting your search box on people's web site, or causing someone's browser to send you search queries first, or paying a phone maker to send you all their search queries, but you have to make enough money from the ads to offset what you pay. And as I mentioned over the last 8 years Google has been paying more and more for their traffic (up to $968M last quarter) and very few entrants into the business are going to compete with that. If you already have a platform (like Mozilla has Firefox, Apple has the iPhone, Facebook has pretty much everyone's Facebook page) so you "own" the ingress point, you can leverage that with a good search engine to make a lot of revenue. But if you need to pay for access to the ingress point, and pay a big chunk to the ad provider, it is really hard to support a lot of infrastructure (which is proportionally expensive to index size). That is the constraint box of search today.

The interesting thing for me is that every quarter, of the last 16, Bing has been making more money per click and Google less, that cost equation is balancing out. That is going to put a lot of pressure on the non-core parts of Google.

To answer your question, Google succeeded well when capturing the value of linkage data to extract page relevance (the original Page Rank patent), they created an advertising incentive which made their algorithm break (you want a billion in-links to your page, no problem! say the black hat SEO folks). Google is still making tons of money on search but you can look at their performance over the last 4 years to see the air is coming out of the balloon. What comes next is still an open question.

Re: Data is at the heart of search, but who has access to it?

#70
post #66

Earlier quoted context omitted.

Chuck, while blekko is a great search engine(especially due to custom search), it's clear that it is very different quality wise from Google.Same for Bing - it's not upto Google.And not for the lack of trying or money(bing). So how do you think Google is succeeding so well, if it's not click stream data? and why can't it be maybe a combination of things that strongly depends on click stream data that others couldn't…

Actually if you do double blind tests you will find that Bing and Google are indistinguishable. We did this at Blekko earlier with our "3 card monte" gambit where you did a query, got back blekko, bing and google results, and got to pick the one with the "best" results for your query. Blekko usually won if it was query we had a slashtag for or if it was a "highly contested" query (lots of ad spend like "no fee credit…

I participated in a blind test between Google, Bing and Yahoo in my Information Retrieval class at a university, back in 2013. The results were: 1) Google, 2) Bing, 3) Yahoo - for every standard IR metric we thought of, which included NDCG@{1, 5, 10}, MRR, MAP.
Post reply on HN