Live data from Hacker News

Mathematics at Google

research.google.com

21–30 of 51 posts

Re: Mathematics at Google

#21
post #3

Look at the slide entitled Gmail (5), and compare the picture with the first graph on my blog post http://jeremykun.wordpress.com/2011/08/11/the-perceptron-and... It just goes to show, Google steals content without attribution just like everyone else.

I am trying really hard to ignore the Futurama picture. Really hard!

At least I gave the source :) And I don't claim I'm any better.

Re: Mathematics at Google

#22
post #16
post #14

Earlier quoted context omitted.

I imagine they have a release process to get something on research.google.com, though. So the company is endorsing it.

What do you expect them to do? Somehow do an image-similarity search throughout the web for every image used in papers by their employees?

They do (almost) the same thing at youtube, don't they?

Re: Mathematics at Google

#23

My mathematician friend pointed out that all that "research at google" requires "experience with large data sets and quantitative analysis". They want statisticians, not mathematicians.

I'm not a googler but I do know a bit of Math and Statistics and many of the approaches they take can be cleverly reduced to optimization problems or matrix math both of which are more the domain of applied mathematicians then statisticians.

Re: Mathematics at Google

#24
post #16
post #14

Earlier quoted context omitted.

I imagine they have a release process to get something on research.google.com, though. So the company is endorsing it.

What do you expect them to do? Somehow do an image-similarity search throughout the web for every image used in papers by their employees?

No, but you would expect some kind of seriousness... Universities don't need to run image-similarity through papers published by their students right? Plagiarism is a big thing.

Re: Mathematics at Google

#25
post #16
post #14

Earlier quoted context omitted.

I imagine they have a release process to get something on research.google.com, though. So the company is endorsing it.

What do you expect them to do? Somehow do an image-similarity search throughout the web for every image used in papers by their employees?

YouTube already has copyright infringement bots that mark copyrighted content and either give the holder the option to put ads on it or take it down.

Search also puts a lot of effort into identifying duplicated content to punish content farms. I've heard they're making progress on detecting algorithmically spun articles too (http://en.wikipedia.org/wiki/Article_spinning).

Oh, and they have it for images. Here's the search result for the one they used: https://www.google.com/search?tbs=sbi:AMhZZivFQlmjC8rcxWC0MZ...

So yes, it'd be trivial for them to do so.

Re: Mathematics at Google

#26
post #8
post #5

Earlier quoted context omitted.

PageRank is definitely still used, although it's doubtfully the sole determinant of ranking results. See http://jeremykun.wordpress.com/2011/06/21/googles-page-rank-...

I agree with your statement, but I'm confused by the story in the link... I haven't come across a message board that doesn't use rel="nofollow" in a few years (that NYT article is from 2010). How would these negative reviews bolster his PageRank?

Google seems to still weight those links, just differently or with less weight.

Re: Mathematics at Google

#27
post #15

Seeing PageRank discussed reminds me of a piece of fun trivia. The idea for PageRank came out of the success of the Science Citation Index, which ranks papers according to how often they have been cited. The idea of trying to study the structure of citations in academia came out of people who were inspired by a 1948 essay, As We May Think . But that essay's main topic was an imagined technology called memex, to be im…

Vannevar Bush's ideas about information organization and consumption in the future were eerily accurate. Reading about the history of Memex and the roots of the Information Architecture field in general is something I highly recommend for anyone interested in Information Science, etc.

Re: Mathematics at Google

#28
post #4

Do they still heavily rely on PageRank? With the amount of traffic data Google has, I would expect more statistical approaches based on what users click (rather than graph algorithms based on how the web is linked) to be the backbone for ranking their results.

It's a closely guarded secret but at the very least we know that every page is annotated with many different signals, which are combined with magical secret sauce.

Re: Mathematics at Google

#29
post #19
post #16

Earlier quoted context omitted.

What do you expect them to do? Somehow do an image-similarity search throughout the web for every image used in papers by their employees?

I would expect them to read it once, and bounce it back to the author saying something along the lines of "Cite your sources, please." That's certainly within Google's power, no?

But if it's on the web, doesn't it belong to Google now?

Re: Mathematics at Google

#30
post #7
post #3

Look at the slide entitled Gmail (5), and compare the picture with the first graph on my blog post http://jeremykun.wordpress.com/2011/08/11/the-perceptron-and... It just goes to show, Google steals content without attribution just like everyone else.

It's one guy who happens to work for Google, rather than the whole company.

Using the same logic, a company can never be said to "do" anything. Everything is (ultimately) done by a real person.
Post reply on HN