Live data from Hacker News

The Long Tail of the English Language

blog.wordsapi.com

1–10 of 55 posts

Re: The Long Tail of the English Language

#2
http://en.wikipedia.org/wiki/Zipf%27s_law

EDIT ADD: Zip's Law was not mentioned directly in the blog post but the reply clarified that the API returns a zipf score. However, a word's zipf ranking is dependent on the corpus used. The Wordapi "About" page[1] says most data came from Princeton WordNet but a sibling comment says it came from a subtitles compilation. If the project could clarify the data sources, it would be helpful.

[1]https://www.wordsapi.com/about

Re: The Long Tail of the English Language

#3
post #2

http://en.wikipedia.org/wiki/Zipf%27s_law EDIT ADD: Zip's Law was not mentioned directly in the blog post but the reply clarified that the API returns a zipf score. However, a word's zipf ranking is dependent on the corpus used . The Wordapi "About" page[1] says most data came from Princeton WordNet but a sibling comment says it came from a subtitles compilation. If the project could clarify the data sources, it woul…

The "frequency" score returned by Words API is the a Zipf score for the word. Ranges from ~1.6 to ~7.6.

Regarding your update - I'll update the About page.

Re: The Long Tail of the English Language

#4
The source material for this frequency count comes from Open Subtitles [http://www.opensubtitles.org/en/search]. Hence the frequencies here apply to spoken English, not written English. In written English the three most common words are "the", "of" and "and", whereas here they are "you", "I", and "the".

Re: The Long Tail of the English Language

#5
I wonder how other languages compare to English. I know English is far from pure.

"The problem with defending the purity of the English language is that English is about as pure as a cribhouse whore. We don't just borrow words; on occasion, English has pursued other languages down alleyways to beat them unconscious and rifle their pockets for new vocabulary." -- James Nicoll

Re: The Long Tail of the English Language

#6

I wonder how other languages compare to English. I know English is far from pure. "The problem with defending the purity of the English language is that English is about as pure as a cribhouse whore. We don't just borrow words; on occasion, English has pursued other languages down alleyways to beat them unconscious and rifle their pockets for new vocabulary." -- James Nicoll

Subtlex has done word frequency counts in a number of languages:

Dutch - http://crr.ugent.be/programs-data/subtitle-frequencies/subtl...

Chinese - http://www.ncbi.nlm.nih.gov/pmc/articles/PMC2880003/

Greek - http://www.bcbl.eu/subtlex-gr/

There's a few others (Polish, French, etc) but I can't find the links for some reason.

Re: The Long Tail of the English Language

#8

Very nice product! As a sidenote: anyway knows a service that provides human associations? Eg What do you associate with "Hacker" as a person "Computer", "Night", "Internet"

It seems like a good idea, but to automate it you'd need to maybe scrap websites and create a count of words that appear in the same sentence. Or you could get really crazy and start comparing subject/object relationships, etc.

Re: The Long Tail of the English Language

#10
post #9

So I'm wondering, if you just learned the 200 most popular words, you might get pretty far in learning a new language, no?

Probably depends on the language. A three-year old supposedly knows about 1,000 English words.

You could probably at least get around.

Post reply on HN