Live data from Hacker News

Suspicious data pattern in recent Venezuelan election

statmodeling.stat.columbia.edu

381–390 of 536 posts

Re: Suspicious data pattern in recent Venezuelan election

#381

Earlier quoted context omitted.

Reporting just the percentages makes sense. Reporting rounded versions of those percentages not only makes sense, but is the universal idiom for reporting percentages. But reporting synthesized vote counts from the percentages --- even from non-rounded percentages --- is not normal. People on this thread are hung up on the reported percentages, but those don't matter in this analysis at all. They're not the problem.…

Maybe I don't understand what you have identified as the problem. My understanding of the article is that the raw tallies should not correspond to "precise" rounded percentages. The article in an addendum points out one way that could legitimately occur (some underling has the totals and rounded percentages but needs the raw tallies and naively multiplies to get them).

I'm summarizing that PPS in my comment. The exculpatory scenario is: (1) start with real numbers, (2) compute percentages, (3) round percentages, (4) discard original numbers, (5) compute new numbers from the round percentages.

Steps (4) and (5) don't have any valid explanation, and few (though maybe some) plausible human error explanations.

As long as we're on the same page that nobody ever had any business reporting the numbers in step (5) --- they're completely fictitious! --- I don't have much to argue about here. The politics aren't interesting to me.

Re: Suspicious data pattern in recent Venezuelan election

#382

This reminds me of the story how the height of Mt. Everest was first measured at exactly 29,000 feet. The surveyor’s boss at the time added two feet thinking no one would trust such an exact number.

I wonder why he added two feet instead of three, or some decimal.

Re: Suspicious data pattern in recent Venezuelan election

#383

Earlier quoted context omitted.

Not difficult at all. Just pick the approximate numbers you want and then introduce a random error of a few percent. (Normal, uniform, doesn't really matter). This is also not hard for statistics experts to detect, but it's much harder to prove (aka you've got plausible deniability). One wonders why they didn't even bother to do fraud slightly better.

Unless done carefully this will almost certainly fail Benford’s Law. Manipulating statistics is harder than you think.

[deleted]

Re: Suspicious data pattern in recent Venezuelan election

#384
post #281

Earlier quoted context omitted.

It was just planned in the US and executed with logistical support from the US with a green light given by Trump (who announced he had no "direct" role in the operation).

That’s crazy you were in the room for that decision. Can you tell us more?

...you didnt read the article at all did you?

Re: Suspicious data pattern in recent Venezuelan election

#385

Earlier quoted context omitted.

Yes, the percentages are the most important number, the number everyone is interested in. The next most important number is the voter turnout. You can verify this by looking at the newspaper headlines of any election. Again unless you are an electioneer no one cares about the raw numbers so it would not be surprising that only the percentages and total are communicated to the public relations department. I don't unde…

> Yes, the percentages are the most important number, the number everyone is interested in. Then why did the hypothetical sub-sub-librarian who put together the final spreadsheet feel the need to go back and repopulate those numbers? Clearly they thought people would want to see them, right? > The next most important number is the voter turnout. You can verify this by looking at the newspaper headlines of any electio…

My recollection is that turnout is usually quoted in both percentage and absolute numbers but quoting it as a percentage requires external data (population demographics) which presumably isn't in the electioneering department.

Why do you have such a hard time believing that election results (e.g. for a union, for school president etc) might be communicated like "55 to 45, 3000 people voted"?

Re: Suspicious data pattern in recent Venezuelan election

#386
post #367

Earlier quoted context omitted.

No. The percentages reported don't matter. These are the same absolute vote numbers; the percentage they work out to is what matters. This seems to be tripping a lot of people up.

It certainly does! Watch the video. No total number is ever reported This means that "total votes" number was worked backwards by a third party. It's not even the CNE's fuckup

No. Watch the video. After Gonzalez' result, he reads the results para otros candidatos†. It's the same as the number in this post. It literally doesn't matter what else happens after this; you can reconstruct the result, after less than 2 minutes of the video's runtime. It's cooked.

as you can see, i am not a fluent speaker of Spanish, and i managed to work it out :)

Re: Suspicious data pattern in recent Venezuelan election

#387

Earlier quoted context omitted.

You can tell a story about a process that publishes these numbers in good faith, but not a story in which the vote counts reported are anything other than fictitious. It is not, in fact, an ordinary sequence of events to take true counts, work out their percentages, round them, discard the original counts , and work back new counts from the rounded percentages. Those new counts are a lie, no matter what the process w…

The entire issue is whether it was a "good faith" mistake. If you concede that...

I don't really think it's even a plausible mistake. Remember, to make the mistake, you have to retain the original raw vote total, but discard the original raw per-candidate totals, then recompute them from the rounded percentages and the retained original total. That doesn't make sense, for reasons having nothing to do with my level of trust in the election authority. It's several steps of extra effort for a result, read live on television by the election authority, that instantly destroys the credibility of the election.

Re: Suspicious data pattern in recent Venezuelan election

#388

Earlier quoted context omitted.

> Yes, the percentages are the most important number, the number everyone is interested in. Then why did the hypothetical sub-sub-librarian who put together the final spreadsheet feel the need to go back and repopulate those numbers? Clearly they thought people would want to see them, right? > The next most important number is the voter turnout. You can verify this by looking at the newspaper headlines of any electio…

My recollection is that turnout is usually quoted in both percentage and absolute numbers but quoting it as a percentage requires external data (population demographics) which presumably isn't in the electioneering department. Why do you have such a hard time believing that election results (e.g. for a union, for school president etc) might be communicated like "55 to 45, 3000 people voted"?

Because I've literally never seen percentages without tallies reported in any context. It's apparently so uncommon that your hypothetical person who created these clearly-not-real numbers felt the need to go backfill them.

Explain that. If it's so unnecessary to report the tallies and people only want to hear the percentages, why did your hypothetical person go back and backfill them?

Re: Suspicious data pattern in recent Venezuelan election

#389

Earlier quoted context omitted.

The entire issue is whether it was a "good faith" mistake. If you concede that...

I don't really think it's even a plausible mistake. Remember, to make the mistake, you have to retain the original raw vote total , but discard the original raw per-candidate totals , then recompute them from the rounded percentages and the retained original total. That doesn't make sense, for reasons having nothing to do with my level of trust in the election authority. It's several steps of extra effort for a resul…

I explained in another comment that these are the two most important numbers in an election and the two numbers everyone cites. Moreover, you might naively assume that they capture all the information about the election because, especially if you are not a STEM major and maybe even then, you might think you can just multiply the numbers to get the per-candidate tallies. And actually for most purposes it might be fine, you rarely need to know the tally to a tenth of a percent accuracy.

Re: Suspicious data pattern in recent Venezuelan election

#390

Earlier quoted context omitted.

I’m not a statistician so I may be confusing it with Zipf’s law. But IIRC tallies from individual precincts should roughly conform to Benford’s law.

I think the concern is that precinct size tends to cluster in ways that mean results can cluster in ways that - for a large portion of the data - does not span a full order of magnitude.

To elaborate, if we imagine a polity with precincts that turn out 10,000 people each election, with two major party candidates that each get between 20% and 80% of the vote, we'd see precisely 0% precincts with a leading digit of 1, much less the ~30% predicted by Benford's law. Of course that doesn't exactly describe any real polity, but it doesn't seem surprising that real elections would be enough like that to screw with the pattern.
Post reply on HN