Live data from Hacker News

A Software Engineer’s Adventures in Learning Mathematics

medium.com

171–180 of 227 posts

Re: A Software Engineer’s Adventures in Learning Mathematics

#171
Part I -- General

> Why not quit my job and go back to school? Well, that’s not really for me. Reading books at my own pace lets me try a subject out without fully committing to it and making it a necessity that I find work based off of it.

I respond here in three parts, the first general, the second more specific and about linear algebra, and the third about statistics.

Really, in school, you still have to learn the stuff. And, there a class might help but won't be enough; so, right, again, during the course and in the hours not in the class, you still have to study and learn the stuff. Or for the class, "mathematics is not a spectator sport". That means you still have to study the stuff, get it between your ears, understand it.

Eventually you can conclude that mostly college and its courses are more for certification than education. For the education, that's heavily up to you.

But there are some dangers in self learning: Not all the ideas and learning materials are good; the really good ones are only a small fraction of all the ones you will likely encounter.

So, in praise of college and profs, they can (1) get you on a good track with good ideas and learning materials and (2) get you unstuck and keep you on track. In such study, it's possible to do too much or too little -- from some experts you can see about what is the right amount to do.

I said "college"; revise that to read "one of the world's best research universities", say, in the US, 1-2 dozen. The ideas and materials you will see in such a university really do stand to be better than nearly all you will encounter otherwise. Such learning is where there's not much substitute for quality. But, still, just to learn the material, you don't really have to enroll in such a university and, instead, just borrow from their course descriptions and materials. Indeed, some of the best such universities are working hard to make their learning materials available to all for free over the Internet. Why? Because those universities want to concentrate on pushing forward with research.

Here's a way: Show up at such a university and appropriate department, in your case, applied math, mathematical sciences, operations research, statistics, whatever. Maybe show up at some public department seminars. Talk with some of the students. Say you have a career going, for your career are interested in what the department is doing, and want to learn more about the program and the mathematical content. So, get some of the students to talk and explain.

Then for some of the courses you are interested in, see who the profs are and look at the course materials, say, texts, handouts, on-line files, etc.

Then after looking at the materials and, say, have made some progress with them, try to get 15 minutes to chat with a prof.

Then take what you just got from that department for free -- broad directions, what they regard as more/less important, texts, course materials, etc. -- and go off and study on your own. When you think that maybe you have some course studied well, try to get a copy of the course final exam or Ph.D. qualifying exam, work through it, and, if it appears you did well, ask a prof to check your solution to a few of the most difficult questions. If your solutions look good, then you will start to look good and may get asked if you would like to apply as a student in the department. So, here you and the prof and department are interviewing each other.

Such things worked for me: (A) In my career I kept running into the work of John Tukey at Princeton and Bell Labs. So, it was stepwise regression, exploratory data analysis, power spectral estimation, convergence and uniformity in topology, his statement equivalent to the axiom of choice, etc. So, I wrote him at Princeton asking about graduate study, mentioned those topics, and got back a nice letter from the department Chair basically inviting me to apply.

(B) I applied to Cornell and got rejected. But largely independently I visited a prof there to discuss optimization, asked about being a grad student, and soon got another letter accepting me to grad study.

(C) At least at one time, the Web site of the Princeton math department said, "Graduate courses are introductions to research by experts in their fields. No courses are given for preparation for the qualifying exams. Students are expected to prepare for the qualifying exams on their own." or some such.

Lesson: In grad school at Princeton, are still expected to learn the qualifying exam materials on your own. Well, to do that, don't have to be in the high rent area around Princeton, NJ.

When I did go to grad school and got my Ph.D., what really saved my tail feathers was what I had done and did on my own as independent study. Some of the courses helped in providing high quality directions and materials, but the real work was nearly all independent.

And, one of the crucial inflection points was when I took a problem in a course but not solved in the course, did some research, and found a solution. The solution was novel, and word spread around the department quickly. My halo got a high polish, and that greatly eased my path through my Ph.D. That is, I had proven results on the most important research academic and Ph.D. bottom line -- I'd done good, novel, "new, correct, significant" research. Then my hair cut or lack there of, sloppy hand writing, occasional upchuck at some bad course material, etc. no longer mattered.

My research looked publishable, and it was -- I did publish it later, in one of the best journals, easily, no revisions. So, lesson: That inflection point was from independent work.

So, don't feel that your independent work is inferior to enrolling as a student. Instead, in the best universities, as a student, nearly all the work is for you to do independently anyway.

Still, I'd repeat -- try to pick the brains, for free, without hurting your present career, of the courses, profs, and materials at a world class department in a world class research university.

Why a research university? World class research is a very high bar and some of the best evidence of good expertise, insight, and judgment in the field -- and you don't want the opposites.

Re: A Software Engineer’s Adventures in Learning Mathematics

#172

Part I -- General > Why not quit my job and go back to school? Well, that’s not really for me. Reading books at my own pace lets me try a subject out without fully committing to it and making it a necessity that I find work based off of it. I respond here in three parts, the first general, the second more specific and about linear algebra, and the third about statistics. Really, in school, you still have to learn the…

Part II -- More Specific

After freshman calculus, it would be good to start with abstract algebra. There you will get handy with (relatively simple versions of, and, thus, a good place to start -- and, I have to say, put you ahead of a surprisingly large fraction of the best chaired professors of computer science) theorems, proofs, sets, axioms, groups (once I published a paper in multi-variate, distribution-free statistical hypothesis tests where the core of the math used some group theory), rings, fields, the integers, rationals, reals, and complex numbers and their leading properties, some important algorithms, e.g., Euclidean greatest common divisor (also the way to find multiplicative inverses in the finite field of the integers modulo a prime number and, thus, the core of a super cute way to do numerical matrix inversion exactly using only short precision arithmetic), number theory and prime numbers (crucial in cryptography), vector spaces (the core of multi-variate statistics and more), and some of the classic results. Some of this material is finite mathematics at times of high interest in computing -- e.g., error correcting coding. For such a course, a good teaching math department would be good. Have a good prof read and correct your early efforts at writing proofs -- could help you a lot.

But starting with linear algebra is also good. There are lots of good books. The grand classic, best as a second text, is P. Halmos, Finite Dimensional Vector Spaces. He wrote this in 1942 after he had gotten his Ph.D. from J. Doob (long the best guy in the US in stochastic processes) and was an assistant to von Neumann at the Institute for Advanced Study (for much of the 20th century a good candidate for the best mathematician in the world). The book is really a finite dimensional introduction to von Neumann's Hilbert space theory.

In

http://www-history.mcs.st-and.ac.uk/BigPictures/Ulam_Feynman...

von Neumann is the guy on the right. The guy on the left is S. Ulam (has a cute result the French mathematician LeCam once called tightness I used once). The guy in the middle is just a physicist! Of course, in that picture they were working up ways to save 1 million US casualties in the Pacific in WWII and were astoundingly successful. Ulam is best known for the Teller-Ulam configuration which in its first test yielded an energy of 15 million tons of TNT. There are rumors that von Neumann worked out the geometry for the US W-88, 475 kilotons in a small package as at

http://en.wikipedia.org/wiki/W88

Von Neumann also has a really nice book on quantum mechanics the first half of which is a totally sweetheart introduction to linear algebra.

Of course, Ulam was an early user of Monte Carlo simulation, still important.

Other linear algebra authors include G. Strang, E. Nering, Hoffman and Kunze, R. Bellman, B. Noble, R. Horn. Also for numerical linear algebra, e.g., G. Forsythe and C. Moler, the LINPACK materials, etc. There are free, on-line PDF versions for some of these. Since the subject has not changed much since Halmos in 1942, don't necessarily need the latest paper copy at $100+!

Re: A Software Engineer’s Adventures in Learning Mathematics

#173

Earlier quoted context omitted.

A great many real, live engineers that I know would object to anyone calling themselves an engineer without a lot of study of the big three: statics, dynamics, and thermodynamics; not just general mathematics. It's one of the reasons I don't call myself an engineer. (But I do get to harsh on them a bit about not being professional programmers---their code isn't pretty.) The other reason, of course, is that the "softw…

Statics, dynamics, and thermo would seem to leave out EE's, wouldn't it?

No, most ABET school's require those courses. Also, if you ever try to get your FE/EIT you will need those courses.

Re: A Software Engineer’s Adventures in Learning Mathematics

#174

Part I -- General > Why not quit my job and go back to school? Well, that’s not really for me. Reading books at my own pace lets me try a subject out without fully committing to it and making it a necessity that I find work based off of it. I respond here in three parts, the first general, the second more specific and about linear algebra, and the third about statistics. Really, in school, you still have to learn the…

Part II -- More Specific After freshman calculus, it would be good to start with abstract algebra . There you will get handy with (relatively simple versions of, and, thus, a good place to start -- and, I have to say, put you ahead of a surprisingly large fraction of the best chaired professors of computer science) theorems, proofs, sets, axioms, groups (once I published a paper in multi-variate, distribution-free st…

Part III -- Statistics

For statistics, that is a messy field. It has too many introductory texts that over simplify the subject and not enough well done intermediate or advanced texts.

Also the subject has essentially a lie: They explain that a random variable has a distribution. Right, it does. Then they mention some common distributions, especially Gaussian, exponential, Poisson, multinomial, and uniform. Then the lie: The suggestion is that in practice we collect data and try to find the distribution. Nope: Mostly not. Mostly in practice, we can't find the distribution, not even of one random variable and much less likely for the joint distribution of several random variables (that is, of a vector valued random variable). Or, to estimate the distribution of a vector valued random variable commonly would encounter the curse of dimensionality and require really big big data. Instead, usually we use limit theorems, techniques that don't need the distribution, or in some cases make, say, a Gaussian assumption and get a first-cut approximation.

Early in my career I did a lot in applied statistics but later concluded I'd done a lot of slogging through a muddy swamp of low grade material.

A clean and powerful first cut approach to statistics is just via a good background in probability: With this approach, for statistics, you take some data, regard that as values of some random variables with some useful properties, stuff the data into some computations, and get out data that you regard as the values of some more random variables which are the statistics. The big deal is what properties the output random variables have -- maybe they are unbiased, minimum variance, Gaussian, maximum likelihood, estimates of something, etc.

For this work you will want to know the classic limit theorems of probability theory -- weak and strong laws of large numbers, elementary and advanced (Lindeberg-Feller) versions of the central limit theorem, the law of the iterated logarithm (and its astounding application to an envelope of Brownian motion), and martingales and the martingale convergence theorem ("the most powerful limit theorem in mathematics" -- it's possible to have making applications of that result much of a successful academic career). And, generally beyond the elementary statistics books, you will want to understand sufficient statistics (and the astounding fact that, for the Gaussian, sample mean and variance are sufficient with generalizations to the exponential family) and, also, U-statistics where the order of the input data makes no difference (and order statistics are always sufficient). Sufficient statistics is really from (a classic paper by Halmos and Savage and) the Radon-Nikodym theorem (with a famous, very clever, cute proof by von Neumann), and that result is in, say, the first half of W. Rudin, Real and Complex Analysis (with von Neumann's proof).

Also with the Radon-Nikodym theorem, can quickly do the Hahn decomposition and, then, knock off a very general proof of the Neyman-Pearson result in statistics. How 'bout that!

Thus, to some extent to do well with statistics, both for now and for the future, especially if you want to do some work that is original, you will need much of the rest of a good ugrad major in math and the courses of a Master's in selected topics in pure/applied math.

So, for such study, sure, at one time Harvard's famous Math 55 used the Halmos text above along with W. Rudin, Principles of Mathematical Analysis (calculus done very carefully and a good foundation for more), and Spivak, Calculus on Manifolds, e.g., for people interested in more modern approaches to relativity theory (but Cartan's book is available in English now). It may be that you are not interested in relativity theory or the rest of mathematical physics -- fine, and that can help you set aside some topics.

Then, Royden, Real Analysis and the first half of Rudin's R&CA as above, along with any of several alternatives, cover measure theory and the beginnings of functional analysis. Measure theory does calculus again and in a more powerful way -- in freshman calculus, want to integrate a continuous function defined on a closed interval of finite length, but in measure theory get much more generality.

And measure theory also provides the axiomatic foundation for modern probability theory and of random variables. Seeing that definition of a random variable is a real eye opener, for me a life-changing event: Get a level of understanding of randomness that cuts out and tosses into the dumpster or bit bucket nearly all the elementary and popular (horribly confused) treatments of randomness.

Functional analysis? Well, in linear algebra you get comfortable with vector spaces. So, for positive integer n and the set of real numbers R, you get happy in the n-dimensional vector space R^n. But, also be sure to see the axioms of a vector space where R^n is just the leading example. You want the axioms right away for, say, the (affine) vector subspace of R^n that is the set of all solutions of a system of linear equations. How 'bout that!

Then in functional analysis, you work with functions and where each function is regarded as a point in a vector space. The nicest such vector space is Hilbert space which has an inner product (essentially the same as angle or in probability covariance and in statistics correlation) and gives a metric in which the space is complete -- that is, as in the real numbers but not in the rationals, a sequence that appears to converge really has something to converge to. Then wonder of wonders (really, mostly due just to the Minkowski inequality), the set of all real valued random variables X such that the expectation (measure theory integral) E[X^2] is finite is a Hilbert space, right, is complete. Amazing, but true.

Then in Hilbert space, get to see how to approximate one function by others. So, in particular, get to see how to approximate a random variable don't have by ones you do have -- might call that statistical estimation and would be correct.

Then can drag out the Hahn-Banach result and do projections, that is, least squares, that is, in an important sense (from a classic random variable convergence result you should be sure to learn), best possible linear approximations. And maybe such an approximation is the ad targeting that makes you the most money.

So, that projection is a baby version of regression analysis. There's a problem here: The usual treatments of regression analysis make a long list of assumptions that look essentially impossible in practice to verify or satisfy and, thus, leave one with what look like unjustified applications.

Nope: Just do the derivations yourself with fewer assumptions and get fewer results but still often enough in practice. And they are still solid results.

For the usual text derivations, by assuming so much, they get much more, especially lots of confidence intervals. In practice often you can use those confidence interval results as first-cut, rough measures of goodness of fit or some such.

But the idea of just a projection can give you a lot. In particular there is an easy, sweetheart way around the onerous, hideous, hated over fitting -- it seems silly that having too much data hurts, and it shouldn't hurt and doesn't have to!

And the now popular practice in machine learning of just fitting with learning data and then verifying with test data, with some more considerations which are also appropriate, can also be solid with even fewer assumptions.

Go for it!

Re: A Software Engineer’s Adventures in Learning Mathematics

#175
post #78

This topic comes up every so often. I think my previous comment applies here [1]: I started a Math degree after 16 years of programming without any Math beyond high school (the highest being high school calculus). Most of my work as a software developer didn't require any "higher" Maths. Once I began studying math, including Modern Algebra, Analysis, Graph Theory, Category Theory, etc., I realized I understood many t…

I'm in a similar boat as you. I went up through Calculus in school and hated it. Around 5 years ago I developed a strong interest in learning relevant applied math and have been enjoying it since. Some things I'd add: 1) Math is fun! If you have the aptitude and disposition to enjoy writing software you'll love working out math problems. They're little nuggets of mental stimulation that you can work on with just some…

I wish there were some sort of open-source prerequisite chain of what order to learn any subjects in. That's the hardest part of self-learning. Elon Musk had a reddit AMA recently where he was asked how he knows so much - he responded that he thought everyone had the capability of learning more than they thought, but that the key was to look at knowledge as a semantic tree. If you learn things in the wrong order, they won't have anything to hang off of.

Re: A Software Engineer’s Adventures in Learning Mathematics

#176

Earlier quoted context omitted.

Statics, dynamics, and thermo would seem to leave out EE's, wouldn't it?

No, most ABET school's require those courses. Also, if you ever try to get your FE/EIT you will need those courses.

An intro course would hardly seem to qualify as "a lot of study".

Re: A Software Engineer’s Adventures in Learning Mathematics

#177
post #103

Earlier quoted context omitted.

>I think it's dishonest to call yourself a Software Engineer and not have an engineering degree. Definitely semantics. For example, I have a BS/MS in applied math from a good engineering school university. I am a software engineer mainly working with ECEs, physics, and other math guys, who are all "software engineers". I see where you're coming from about being a SW eng without knowing a lot of math - but there are o…

Right, an engineering degree is a guarantee that you know the important points of what it means to be an engineer. Ethics, engineering process, math, etc. When you get an engineering degree from a certified university, you are an engineer. It's a guarantee that (at one point, at least), you possessed the set of skills and knowledge determined by a board of professionals to be necessary for a career of engineering. Ot…

Since we're playing semantics,

>When you get an engineering degree from a certified university, you are an engineer.

That's not true. When you get a degree you have an engineering degree. When you get hired and employed as an engineer, you are an engineer. You have an engineering degree and I have a mathematics degree. Our employers hire people for engineer positions and call them such. If I were a professor of math, i could call myself a professor. If I were a mathematician at the NSA, a Mathematician. But my employer calls me an Engineer. The degree does not do that.

Re: A Software Engineer’s Adventures in Learning Mathematics

#178

This topic comes up every so often. I think my previous comment applies here [1]: I started a Math degree after 16 years of programming without any Math beyond high school (the highest being high school calculus). Most of my work as a software developer didn't require any "higher" Maths. Once I began studying math, including Modern Algebra, Analysis, Graph Theory, Category Theory, etc., I realized I understood many t…

> The range of problems I could tackle as a programmer was limited by math. It turns out this was partly true. Would you mind expanding on this? I've often thought about going back to school for a math degree, or at least for the core degree courses. Did functional programming, particularly in a pure language like Haskell, become easier to reason about once you had studied Category Theory in depth?

There's no dependency between learning functional programming, including Haskell, and category theory. I say this as someone who is reading a CT book (Conceptual Mathematics) after being first exposed to the topic by learning Haskell.

Re: A Software Engineer’s Adventures in Learning Mathematics

#179

This topic comes up every so often. I think my previous comment applies here [1]: I started a Math degree after 16 years of programming without any Math beyond high school (the highest being high school calculus). Most of my work as a software developer didn't require any "higher" Maths. Once I began studying math, including Modern Algebra, Analysis, Graph Theory, Category Theory, etc., I realized I understood many t…

> i.e. more than one way to skin a cat.

While this is true, I've always found applying mathematical analyses such as algebraic reasoning to computation admits a single, minimal, canonical solution in the end.

Re: A Software Engineer’s Adventures in Learning Mathematics

#180

The comments are amusing. It seems there's a popular opinion that to be an engineer requires a higher understanding of mathematics than what most mere mortals require. My understanding was that to call yourself an engineer one must be capable of building robust and efficient things, whether that thing is a plane, or a bridge, complex software, even a website, matters not, so long as it is robust and efficient. The de…

As an electrical engineer, I like to quip that 99% of the work I do requires only V=IR, or slight variations / derivations thereof.

Of course, it's that remaining 1% that will fuck you if you don't have the maths chops. And this is as a consulting engineer, which arguably uses the least maths. Once you get into actual design or analysis, that number becomes a sliding scale in the other direction very quickly.

Post reply on HN