What about the YC motto "Make something people
want."?
I mean, reading over these requests, my guess is
that, while some of them, if successful, would help
lead to a better life for nearly everyone and save
the world, so far not many people would "want" the
results in the sense needed by a startup.
Some of the requests are nearly hopeless: E.g., the
US Federal Government via DoE, NSF, and NIH have
been spending billions on research in energy and
medicine for decades. The idea that a YC start up
could do a lot better makes most long shots look
like sure things.
Next, a lot of these requests ask for some darned
challenging research projects, and just a research
project is one of the worst insults passed out by
the venture capital community. Instead, no matter
what the Web sites of early stage venture firms
suggest, such firms want to see traction, not
research projects, not even projects to write
software for research already successfully done, not
even to go live with software already written from
research projects already done.
Example? Okay, want a research project to solve a
big problem? Okay, consider security and
reliability of the complex systems of large server
farms and networks. We'd like to do better, right?
For this, the first step is essentially near
real-time monitoring for ASAP detection. So, we
want detectors.
First big problem is getting a good combination of
rates of false alarms and missed detections; I'm
correct here; think for a few minutes and otherwise
trust me on this one. Or, we're trying for a good
combination of false positives and false negatives.
Or for a good combination of Type I and Type II
error.
Right, you guessed it, oh how you guessed it: Such
detection has just two ways to be wrong -- a false
alarm where we say that the system is sick when it
is healthy and a missed detection where we say that
the system is healthy when it is sick. Inescapable.
Have any doubts, then think for two minutes. With
me again now?
Okay: So, right, such monitoring and detection is a
case of ASAP, essentially real-time, statistical
hypothesis testing. I know; I know; you don't want
anything that is just statistical. Neither do I.
Tough stuff. We're necessarily, inescapably stuck-o
never the less. Or, want a detector with no false
alarms? Got that one for you -- just turn off the
detector. Want a detector with no missed
detections? Got one of those, too -- just sound the
alarm all the time. Yes, some detectors
occasionally make correct detections and have
essentially no false alarms, but such detectors will
detect only problems of a very narrow kind and
otherwise have a high rate of missed detections.
Arguing is futile -- the hard stuff isn't here, and
I'm correct here. We're talking research here, like
YC now seems to want to see, and research can be
tough stuff to swallow. Keep reading ....
Now, what the heck to do about this? Okay, we've
got some good news that, right Andreessen Horowitz
should understand quickly: We can get data on each
of several variables, maybe dozens or hundreds, at
data rates of a point each few seconds up to
hundreds of points a second. At a big, complex
system, we're talking big data. So, we want our
statistical hypothesis test to be multi-dimensional.
Sure, go to the library and find a lot of those,
right? Wrong. You won't find much. Next, most of
the material you see on statistical hypothesis tests
wants the probability distribution of the data when
the system is healthy. Tough since for these
complex systems there's no theory that will give you
means of finding such distributions (we're talking
multi-dimensional), and, even with big data,
anything like accurate estimates of
multi-dimensional distributions is agony with the
curse of dimensionality. So, now what? Okay, we
want to be distribution-free, that is, have a
statistical hypothesis test that makes no
assumptions about the probability distribution.
So, how many multi-dimensional, distribution-free
statistical hypothesis tests did you find in the
library? Not a lot. Maybe the only ones you found
were mine. Mine? Yup.
But, when the dust settles, we do get a (large class
of) genuine statistical hypothesis tests that are
both multi-dimensional and distribution-free. So,
right, with meager/standard assumptions, as is
standard we can calculate false alarm rate and set
it in advance and get it exactly in practice. For
detection rate, as is usually the case we don't have
enough data to use the best possible Neyman-Pearson
result, but there is good reason to regard the
detection rate as relatively high.
So, any large server farm or network doing important
work and interested in security and reliability, ,
that is, nearly all of them, will be interested,
right? And any VC firm, too, right? Nope. Don't
hold your breath waiting. I only wrote nearly every
information technology venture firm in the country,
indeed, some months before the bubble burst in 2000.
Responses? Even during the days of big bottles of
Wonder-Bubble, none or f'get about it.
Lesson: Research, even for a big problem at
important enterprises, even done research, even with
algorithms to make the computing fast, even with
prototype software running, even with a research
paper that passed high quality peer review, doesn't
get venture funding. I learned that lesson. Here
HN and YC can learn it now or learn it later. Now
is easier.