Live data from Hacker News

Ranking Some Generals in the History of Warfare

medium.com

21–30 of 41 posts

Re: Ranking Some Generals in the History of Warfare

#21
I imagine the average ability modern generals receive is due to far more accurate troop count measurements. That also explains the Western focus, as Wikipedia often doesn't have troop counts for historic battles elsewhere. Probably helps Caesar and Napoleon as well, both excelled at propaganda about their feats.

Re: Ranking Some Generals in the History of Warfare

#22
post #3

"Ranking Almost Every General in the History of Warfare" Not even close to "almost every general". Even taking into consideration his own admission that he left out the Mongols, his analysis is obviously very, very Western centric. Somehow the omission of the most populous continent in the world and probably throughout history didn't dissuade the author from making such a bold claim. The most obvious omissions were t…

Garbage in, garbage out. If your data set comes from scraping wikipedia, you're going to have these kinds of flaws and omissions. What's the alternative though, other than to hire an army of grad students? If flawed, at least it's interesting. I wish the author had gone into more detail about counterintuitive results (like Rommel), but part of the point of an exercise like this is to find instances where the model di…

Eh, scraping any encyclopedia would introduce insurmountable bias and error. As you say, only primary sources are usable here.

Re: Ranking Some Generals in the History of Warfare

#23
post #4
post #3

"Ranking Almost Every General in the History of Warfare" Not even close to "almost every general". Even taking into consideration his own admission that he left out the Mongols, his analysis is obviously very, very Western centric. Somehow the omission of the most populous continent in the world and probably throughout history didn't dissuade the author from making such a bold claim. The most obvious omissions were t…

OK, we'll put "some" above.

Has HN considered color coding title changes, or even the portions of the title changed in smaller or subtle changes like here?

Re: Ranking Some Generals in the History of Warfare

#24
post #3

"Ranking Almost Every General in the History of Warfare" Not even close to "almost every general". Even taking into consideration his own admission that he left out the Mongols, his analysis is obviously very, very Western centric. Somehow the omission of the most populous continent in the world and probably throughout history didn't dissuade the author from making such a bold claim. The most obvious omissions were t…

It's almost 2018 and you're still peddling intolerant anti-sensationalism wrongthink? Get with the times! /s

Re: Ranking Some Generals in the History of Warfare

#25
post #22

Earlier quoted context omitted.

Garbage in, garbage out. If your data set comes from scraping wikipedia, you're going to have these kinds of flaws and omissions. What's the alternative though, other than to hire an army of grad students? If flawed, at least it's interesting. I wish the author had gone into more detail about counterintuitive results (like Rommel), but part of the point of an exercise like this is to find instances where the model di…

Eh, scraping any encyclopedia would introduce insurmountable bias and error. As you say, only primary sources are usable here.

Yeah, didn't mean that as a knock against wikipedia specifically.

Re: Ranking Some Generals in the History of Warfare

#26
WAR? haha I mean... I guess that would lead one to try and apply some analysis to generals.

This seems like a fun exercise in modeling data and doing some quick analysis. Obviously not for doing any worthwhile analysis or grading our existing general officer corps or putting any other weights in there.

For true analysis of tactical prowess you're looking at battle-to-battle analysis I'd imagine and metrics just don't go back that far. "most flanks covered", "most miles gained", "most efficiency per round of ammunition"... rather dark either way.

Well done for the purpose and nice post about how to gather / munge data!

Out of curiosity, anybody know what Schwarzkopf's stats would have been?

Re: Ranking Some Generals in the History of Warfare

#27

Was Borodino really a victory for Napoleon? Wasn't it more of a draw? Sure, the Russians retreated afterwards, so it didn't stop Napoleon's drive on Moscow. Still...

It was a Pyrrhic victory. At the end of battle, the Russian army retreated from the field and the French remained - so it was a victory by definition. But the French army could not replace its casualties or secure its supply lines, so the victory in battle resulted in defeat in the overall war.

Re: Ranking Some Generals in the History of Warfare

#29

WAR? haha I mean... I guess that would lead one to try and apply some analysis to generals. This seems like a fun exercise in modeling data and doing some quick analysis. Obviously not for doing any worthwhile analysis or grading our existing general officer corps or putting any other weights in there. For true analysis of tactical prowess you're looking at battle-to-battle analysis I'd imagine and metrics just don't…

Not sure the intent of the author was to cover operational-level commanders. Even if Gulf War I did count, coalition forces outnumbered the Iraqis in sum total, so Schwarzkopf's WAR probably wouldn't be very good.

Re: Ranking Some Generals in the History of Warfare

#30
I can't seem to comment anymore on the original article, so I'm going to write here in the hope that the author will see it.

Lots of people are commenting about the poor data set selection. I think that's understandable given that the author isn't a historian but rather a baseball geek and data scientist. Although I can't help but feel a little bit of anger to see (speaking as a software engineer and a history Ph.D. dropout) yet another example of a technical person blithely wandering into a field that has been studied for thousands of years and make very grandiose claims without even a cursory study of the field. Buy hey, that's tech for you, always disrupting (I mean that both sarcastically and not sarcastically at the same time).

What I want to address is more fundamental than getting the data set right, however. The author doesn't seem to understand that the very nature of the historical record is highly subjective.

1) Even if Wikipedia nailed every statistic, the statistics themselves about wins and losses, troop numbers, lengths of battle, places, etc. are increasingly unreliable as you go back in time. In some texts and historiographical traditions, the numbers are not just unreliable, they're arguably cut from whole cloth. Anything past, say, 1000 A.D. in Western Europe, for example, is highly disputed. Biases in old texts aside, we have enormous gaps in what texts have survived to this day. History isn't written by the victors, but the dried wood pulp and calf skins that it's written on is selectively preserved by them. The author's model has no awareness of the history of historiography.

(Tangentially, Napoleon, whom the author's model rates as the greatest general of all time, was indirectly responsible for a huge amount of destruction of Europe's archives after issuing orders to transfer archives from across the continent to Paris. Early modern logistics meant that huge portions of documents were destroyed in transport. I remember when I was working in the Vatican Secret Archives, something like 1/3rd of that archive was destroyed, and that's one of the main archives for European and world history.)

2) Even if we had all the numbers exactly perfect, what counts as a victory and what counts as a loss is highly subjective. Was the North Vietnamese Tet Offensive a victory or a loss? For whom? In what sense? These philosophical questions can't be answered by a model, at least not without the philosophical assumptions being made explicit.

This question is fundamentally one that requires a nuanced approach. I think that data-driven approaches can really help, but the author's model needs not only more refinement, it also needs to acknowledge more of the confounding factors involved. I encourage the author to keep working on it.

Post reply on HN