Live data from Hacker News

Estimate your English vocabulary size

testyourvocab.com

51–60 of 331 posts

Re: Estimate your English vocabulary size

#51

Think I have the lowest score here. 16,400 words. English is not my native language but I speak English daily and I wouldn't say my English is bad. Pretty disappointed with the score and also surprised the median is way way higher than I expected. Edit: And, also to add, I followed 2 criteria for whether I know the word or not. 1. What's the absolute definition? 2. And can I find the equivalent or meaning of it in my…

I think a lot of the words are ones that you will only ever encounter in reading fiction (and flower/older fiction at that). Also, keep in mind that this test doesn't include field-specific technical jargon.

Re: Estimate your English vocabulary size

#52
post #37
post #32

lots of people here are saying they scored lower than what they expected, and that maybe other people cheated. that could be it, but it could also be that hacker news folks tend to be overconfident. this would match the stereotype of this group being mainly male nerd entreprenuers, which could score worse on things like this but perceive themselves to score much higher (a feeling not a fact backed by studies that i c…

I got just under the median on this test, but I scored around the 99th percentile in the SAT verbal, and I feel it's reasonable to say the two test approximately the same things. It seems unlikely that I've slipped so far in just two years :) I never really considered Hacker News to be full of overconfident people. If anything, to me being part of HN is a humbling experience. It reminds me that there are so many peop…

I'm pretty much in the same boat. This scored me below median, but I scored in the 99th percentile on the SAT verbal as well (although that was a decade ago for me).

I think one contributing factor is the structure of the questions. If you ask me if I KNOW at least one definition for the word mawkish for example, I'll choose no I don't know a definition of it. I do however have a good enough feel for the word, that I can almost guarantee I can get a SAT style analogy with it correct, or if you gave me a multiple choice selection of definitions I can probably pick out the right one. I don't consider either of those skills equivalent to actually knowing a definition of a word.

Re: Estimate your English vocabulary size

#53
post #42

I'm not sure if vocabulary size matters once you reach around 25,000 words. The words I didn't know were in part because I've never had any need to know them; if I had run into any of them while reading anything written in the past 80 years, I'd be angry at the author for showing off. When I was young, I thought that if I wanted to be a writer I should have a huge vocabulary... but now, when choosing words/synonyms I…

> I'm not sure if vocabulary size matters once you reach around 25,000 words.

This is what I was thinking. I scored 34K, and rarely encounter a word that I don't understand in regular speech or reading. I also know several thousand jargon words, none of which were on that test. I know what I need to know. Memorizing another 16K words to reach Shakespeare's magical 50K (and feel good about myself) would be a waste of precious mental resources.

Re: Estimate your English vocabulary size

#54
post #29

I got 37,300. They claim this is not quite 95th percentile, which I am a tad skeptical accurately represents my vocabulary-size percentile relative to the general population. Perhaps this survey is being forwarded around unusually literate people at the top end, or more than 5% of responders are cheating. Where are the fake words to catch cheaters? I Googled a lot of what I didn't recognize, and everything I checked…

Eliezer, I think you appreciate some of the ideas behind the FAQ I'll repost here with adaptation to the current situation:

VOLUNTARY RESPONSE POLLS

As I commented previously when we had a poll on the ages of HNers, the data can't be relied on to make such an inference. That's because the data are not from a random sample of the relevant population. One professor of statistics, who is a co-author of a highly regarded AP statistics textbook, has tried to popularize the phrase that "voluntary response data are worthless" to go along with the phrase "correlation does not imply causation." Other statistics teachers are gradually picking up this phrase.

-----Original Message----- From: Paul Velleman [SMTPfv2@cornell.edu] Sent: Wednesday, January 14, 1998 5:10 PM To: apstat-l@etc.bc.ca; Kim Robinson Cc: mmbalach@mtu.edu Subject: Re: qualtiative study

Sorry Kim, but it just aint so. Voluntary response data are worthless. One excellent example is the books by Shere Hite. She collected many responses from biased lists with voluntary response and drew conclusions that are roundly contradicted by all responsible studies. She claimed to be doing only qualitative work, but what she got was just plain garbage. Another famous example is the Literary Digest "poll". All you learn from voluntary response is what is said by those who choose to respond. Unless the respondents are a substantially large fraction of the population, they are very likely to be a biased -- possibly a very biased -- subset. Anecdotes tell you nothing at all about the state of the world. They can't be "used only as a description" because they describe nothing but themselves.

http://mathforum.org/kb/thread.jspa?threadID=194473&tsta...

For more on the distinction between statistics and mathematics, see

http://statland.org/MAAFIXED.PDF

and

http://escholarship.org/uc/item/6hb3k0nz

I think Professor Velleman promotes "Voluntary response data are worthless" as a slogan for the same reason an earlier generation of statisticians taught their students the slogan "correlation does not imply causation." That's because common human cognitive errors run strongly in one direction on each issue, so the slogan has take the cognitive error head-on. Of course, a distinct pattern in voluntary responses tells us SOMETHING (maybe about what kind of people come forward to respond), just as a correlation tells us SOMETHING (maybe about a lurking variable correlated with both things we observe), but it doesn't tell us enough to warrant a firm conclusion about facts of the world. The Literary Digest poll

http://historymatters.gmu.edu/d/5168/

http://www.math.uah.edu/stat/data/LiteraryDigest.pdf

is a spectacular historical example of a voluntary response poll with a HUGE sample size and high response rate that didn't give a correct picture of reality at all.

When I have brought up this issue before, some other HNers have replied that there are some statistical tools for correcting for response-bias effects, IF one can obtain a simple random sample of the population of interest and evaluate what kinds of people respond. But we can't do that here on HN, nor can we for the online vocabulary estimation.

Another reply I frequently see when I bring up this issue is that the public relies on voluntary response data all the time to make conclusions about reality. To that I refer careful readers to what Professor Velleman is quoted as saying above (the general public often believes statements that are baloney) and to what Google's director of research, Peter Norvig, says about research conducted with better data,

http://norvig.com/experiment-design.html

that even good data (and Norvig would not generally characterize voluntary response data as good data) can lead to wrong conclusions if there isn't careful thinking behind a study design. Again, human beings have strong predilections to believe certain kinds of wrong data and wrong conclusions. We are not neutral evaluators of data and conclusions, but have predispositions (cognitive illusions) that lead to making mistakes without careful training and thought. Here, the conclusion "those other guys are cheating and that dragged down my vocabulary percentile score" is an example of a conclusion resulting from human predispositions.

Another frequently seen reply is that sometimes a "convenience sample" (this is a common term among statisticians for a sample that can't be counted on to be a random sample) of a population offers just that, convenience, and should not be rejected on that basis alone. But the most thoughtful version of that frequent reply I recently saw did correctly point out that if we know from the get-go that the sample was not done statistically correctly, then even if we are confident (enough) that HN participants are young or that their vocabularies are large, we wouldn't want to extrapolate from that to conclude that the users of any technology site are young, or that respondents to online surveys as a whole have large vocabularies.

On my part, I wildly guess that most HNers are younger than I am in part because this kind of poll recurs often on HN. I similarly guess that participants in online surveys of vocabulary size are likely to have larger vocabularies than average people in the general public because most people I meet find discussions of word meanings boring. But neither guess gives me a good quantitative basis for estimating how much users here differ from the general population.

Re: Estimate your English vocabulary size

#56
I got 28,800 http://testyourvocab.com/?r=38317 So apparently I should be 31 instead of almost 21.

I had a phase around 7th through 10th grade where I thought learning lots of vocabulary would make me smarter, especially words others didn't know well. (And so I'd use them in English essays for Extra Points since your grades are often determined by how little sense you make, because if the reader doesn't understand it obviously it's too smart for them!) I also had a general grammar nazi-ism.

Anyway, I think this exchange kind of tipped me over the edge to stop caring. (Of course that's led to forgetting a lot.)

William Faulkner, on Ernest Hemingway: "He has never been known to use a word that might send a reader to the dictionary."

Hemingway: "Poor Faulkner. Does he really think big emotions come from big words? He thinks I don't know the ten-dollar words. I know them all right. But there are older and simpler and better words, and those are the ones I use."

Of course, having some background in French and Latin probably helps for inferring a few words.

Re: Estimate your English vocabulary size

#57

Don't care too much for percentile and age stats at the end.. clearly doesn't represent general pop. What IS interesting tho, from a language learners perspective, is the vocab size estimation. A metric a lot of us use as a rough benchmark of vocab needed for fluency in a foreign language is 10,000words. Comparing this with what an educated adult native speaker knows in their own language (using my own truthful score…

I agree: the test seems to correctly estimate vocabulary size, but is incorrect in calculating percentile.

I got 12900 (English is my second language) and it seems to be about accurate. I also feel that I'm fluent in English, and it confirms that 10000 words is sufficient for fluency.

The test is also missing professional lingo: where are such words as SQL, lisp, ai, startup, PG, HN and other that we all know so well?

Re: Estimate your English vocabulary size

#59
"You will never become proficient in a foreign language by studying vocabulary lists. Rather, you must hear and speak (or read and write) the language to gain proficiency. The same is true for learning computer languages."

Coincidentally, I just happened to come across this quote in Peter Norvig's "Paradigms of Artificial Intelligence Programming".

Re: Estimate your English vocabulary size

#60
I scored 38500 - seemed to be a test that would be helped by reading a lot of older fantasy literature, where 'terpsichorean' and 'turpitude' (to give a couple 'terp' examples that spring to mind) are the sort of words that authors like Jack Vance liked to wheel out in order to create a mood.

I'm not sure that the people suggesting that the failure to correlate with the SAT adds much; I don't think the SAT really goes all-out of the more flowery bits of archaic vocabulary in the way that this test did.

My 3rd grade son got 10200, and enjoyed discussing the words he didn't get. I think every 3rd grader should know "mawkish". :-)

Post reply on HN