Live data from Hacker News

Larry Page on Real Time Google: We Have To Do It

readwriteweb.com

21–27 of 27 posts

Re: Larry Page on Real Time Google: We Have To Do It

#21
post #20
post #18

Earlier quoted context omitted.

2 thoughts: 1) I worked on the search engine and Larry has been saying this for years. I worked on search quality back in 2005 and even back then he was talking about indexing everything and indexing it in seconds instead of hours. 2) Search is won on the margins. Yahoo and Google do equally well on most queries, but users decide which engine is better based on how it performs over all of the types of searches they h…

Update: Hey, you're the guy who launched Google Transit! I could use your advice ( See http://www.ridecell.com/gt/about/ ). Can I email you? I think indexing everything in seconds could definitely be a competitive advantage. I haven't tried Yahoo in years and back then it did much worse than Google. If it has improved this much, it makes sense that the competition is at the margins. On the other hand, TechCrunch/Twit…

[deleted]

Re: Larry Page on Real Time Google: We Have To Do It

#23

Imagine if this had happened 5-8 years ago and instead of using Twitter/FB/MySpace for the 'activity stream' it used IM statuses from all the major IM networks (ICQ, AOL, MSN, jabber etc.) Bizaro world for sure but interesting to ponder.

I wrote exactly this 4 years ago :)

It was called AwayGrabber (www.awaygrabber.com).

I wrote a overly complex crawler in C to grab away messages from IM networks as fast as rate limiting would allow. Then created a web frontend for viewing all of the status messages from your friends.

It was cool since in most clients at the time you needed to click on a friend and select "get info" for each status you wanted to read. Feeds for status make much more sense. However, I got tired of trying to reverse engineer the changes in various closed protocols (oscar, etc). So I did more than ponder this when I was in college, I tried it.

Re: Larry Page on Real Time Google: We Have To Do It

#24
post #12
post #5

How does Google figure out relevance in realtime? With twitter, it's user driver content with tags etc. But with the net at large, blogs etc, this becomes difficult. Incoming links etc are hard to determine in real time (primarily because they haven't occurred yet).

The site's PR, uptime/age, update rate, uniqueness of content. By this measure HN is near the ideal: it has amazing inbound PR but does not link out much. It's been running fast and fine for ~3 years, the content is often unique, in the sense that it contains phrases Googlebot has never encountered before. An experiment: here is a search that matches an exact phrase in this comment. http://www.google.com/search?q=%22…

12 hours later, still no match.

Re: Larry Page on Real Time Google: We Have To Do It

#25
post #8

When I watched the LOST season finale last Wednesday I was trying to find out the answer to the question "What lies in the shadow of the statue" right after the show ended. (During the show, the answer was spoken in a different language) Google real-time search had picked up the answer from both a TV forum and Yahoo! Answers in about 5 minutes and it took Twitter about 55 minutes before anyone had an answer in the se…

All hail John Locke.

Re: Larry Page on Real Time Google: We Have To Do It

#26
post #12
post #5

How does Google figure out relevance in realtime? With twitter, it's user driver content with tags etc. But with the net at large, blogs etc, this becomes difficult. Incoming links etc are hard to determine in real time (primarily because they haven't occurred yet).

The site's PR, uptime/age, update rate, uniqueness of content. By this measure HN is near the ideal: it has amazing inbound PR but does not link out much. It's been running fast and fine for ~3 years, the content is often unique, in the sense that it contains phrases Googlebot has never encountered before. An experiment: here is a search that matches an exact phrase in this comment. http://www.google.com/search?q=%22…

The quotes seem to be the problem, works without for me.

Re: Larry Page on Real Time Google: We Have To Do It

#27
post #24
post #12

Earlier quoted context omitted.

The site's PR, uptime/age, update rate, uniqueness of content. By this measure HN is near the ideal: it has amazing inbound PR but does not link out much. It's been running fast and fine for ~3 years, the content is often unique, in the sense that it contains phrases Googlebot has never encountered before. An experiment: here is a search that matches an exact phrase in this comment. http://www.google.com/search?q=%22…

12 hours later, still no match.

It has appeared in the index after somewhere between 12 and 17 hours.
Post reply on HN