Live data from Hacker News

Big Data and the Topologist (2012)

ldtopology.wordpress.com

21–22 of 22 posts

Re: Big Data and the Topologist (2012)

#21

I tried analysing data using persistent homology. What is not obvious, although they do admit it in one line of every paper, is that is it susceptible to noise :( So it has to go in the bin even though I really want to know what my manifolds look like!

Well, the thing about noise, from a topological point of view, is that persistent homology simply cannot decide whether something is noise or not --- at least not right away. To be more precise: A lot of what PH does is actually some sort of multi-scale Betti number calculation. The Betti numbers count the number of k-dimensional "holes" in a data set. Their calculation is usually done by something that is called "si…

Yeah I understand all that. I sent you an email.

"Now, the problem about real-world data is that it does not come in the form of a simplicial complex." Not really true. real world data comes from a low dimensional manifold complex + noise.

I thought PH would extract the manifold but I could not get a tractable solution.

Re: Big Data and the Topologist (2012)

#22
post #7

I tried analysing data using persistent homology. What is not obvious, although they do admit it in one line of every paper, is that is it susceptible to noise :( So it has to go in the bin even though I really want to know what my manifolds look like!

Just a few days ago the following paper was published on the ArXiv, it specifically tackles the problem of noise. I haven't had time to read it thoroughly, but it also contains a nice introduction to persistent homology for statisticians, who are usually not familiar with the concepts of algebraic topology. http://arxiv.org/pdf/1303.7117v1.pdf

Yeah that looks like a step in the right direction, but the use of purely synthetic examples leaves me a little sceptical it will work in practice. I don't have time to try every papers algorithms out. If they don;t put in real examples I think perhaps they tried it on real data but it still did not work. Persistent Homology is in its infancy, its developed by pure mathematicians who are a bit innocent to the horrors of real data (or willfully ignore pragmatic trivialities in preference to interesting theoretical results).
Post reply on HN