Live data from Hacker News

Researchers: It takes 1.5 MB of data to store language information

medicalxpress.com

41–50 of 59 posts

Re: Researchers: It takes 1.5 MB of data to store language information

#41
post #38
post #35

Earlier quoted context omitted.

Not sure I follow. If (big if) there were really a measureable advantage to having us "pick the direct word quickly, instead of the slower way of deriving it from a rule" , it doesn't follow that irregularity makes that easier. I could memorize "goed" just as easily as I can memorize "went".

I took it the other way around: it's easy to memorize "went" because you use it all the time. If on the other hand a much less common verb like "to satiate" had a very irregular conjugation then it would regularize pretty quickly because nobody but ultra-pedants would bother to remember the exception. I think a decent real example of that is fiancé/fiancée, those are french borrowings and have, at least originally, k…

>I took it the other way around: it's easy to memorize "went" because you use it all the time.

That still wouldn't explain the why of having it like "went" vs "goed".

Sure, it's easy to memorize because we use it all the time, but why have it to memorize it in the first place, versus something like "goed".

So, this theory (I tried to convey above) said that it being irregular placed ensured we don't slow down try to derive it from regular rules, but instead have fast access to a memorized form.

Couldn't we just memorize "goed"? If it's just "frequency of use" that mattered, "went" and "goed" would work just as well.

But the extra idea is that "goed", being regular, would be too easy for us to confuse with thousand of other regular verbs, and not use our "fast recall" mechanism, regardless of that verb being needed all the time.

Not sure if correct - read it years ago. This seems to be related to that:

https://en.wikipedia.org/wiki/Regular_and_irregular_verbs#Li...

Re: Researchers: It takes 1.5 MB of data to store language information

#42
post #24

Earlier quoted context omitted.

> This describes how to summarize a language in a computer system. It describes how to summarize a language from a information theory standpoint (which is something different). There's a good reason to believe evolution would use something close to the most information theoretic efficient implementation (and even if not, it's a good lower bound)

> There's a good reason to believe evolution would use something close to the most information theoretic efficient implementation Do you have a citation? I would be very interested to see it. I hate to make analogies between computing and biological systems, but I immediately think of denormalization in databases, error correcting codes, and database replication as examples of optimizations that take you farther from…

Shannon observed this in his work towards information theory. Humans are able to predict next letters in a stream of natural language with a probability approaching optimal. The original paper is here:

https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf

The summary at wikipedia:

https://en.wikipedia.org/wiki/Entropy_(information_theory)#E...

Notably, PPM, the best compression algorithm (at the time), was not able to predict as well. So that's a quantitative demonstration of optimality in language performance that includes a faculty for understanding syntax and semantics and presumably is linked to phonology, pragmatics etc.

Chomsky discusses optimality in a 1998 talk here:

https://youtu.be/7Sw15-vSY8E, at about 1h 7m an onwards.

He says language evolved rather quickly for such a significant function (estimating it arose quickly ~200kya). He says though many things in evolution are pretty messy, language is close to optimal design. He doesn't provide references, but I think (given the time and academic associations; anyone know?) he may be referring to Optimality Theory:

https://en.wikipedia.org/wiki/Optimality_Theory

Separately, our knowledge of anatomy of the language faculty is that it is highly localized in the Wernicke and Broca areas of the brain. This also suggests a small modification of an existing brain structure.

Additionally, optimality in biological systems has been a productive research direction in a few areas: - Optimality Principles in Biology - https://link.springer.com/book/10.1007/978-1-4899-6419-9 - Near optimal energy use in the metabolism of cell division - https://arxiv.org/pdf/1209.1179.pdf - https://www.quantamagazine.org/the-math-that-tells-cells-wha...

Re: Researchers: It takes 1.5 MB of data to store language information

#43
post #24

Earlier quoted context omitted.

> This describes how to summarize a language in a computer system. It describes how to summarize a language from a information theory standpoint (which is something different). There's a good reason to believe evolution would use something close to the most information theoretic efficient implementation (and even if not, it's a good lower bound)

> There's a good reason to believe evolution would use something close to the most information theoretic efficient implementation. People used to think that about genomes, until we got to know in some detail what's in them.

Explain.

Re: Researchers: It takes 1.5 MB of data to store language information

#44
post #41
post #38

Earlier quoted context omitted.

I took it the other way around: it's easy to memorize "went" because you use it all the time. If on the other hand a much less common verb like "to satiate" had a very irregular conjugation then it would regularize pretty quickly because nobody but ultra-pedants would bother to remember the exception. I think a decent real example of that is fiancé/fiancée, those are french borrowings and have, at least originally, k…

> I took it the other way around: it's easy to memorize "went" because you use it all the time. That still wouldn't explain the why of having it like "went" vs "goed". Sure, it's easy to memorize because we use it all the time, but why have it to memorize it in the first place, versus something like "goed". So, this theory (I tried to convey above) said that it being irregular placed ensured we don't slow down try to…

Your point is interesting but I think you're falling for the same type of fallacies people often have regarding evolution (if you keep going into water for a long time over many generations you'll eventually grow gills!).

Natural languages are not designed, they evolve. The irregular nature of the conjugation of "to go" might just be a remnant of some archaic form and nothing else, in the same way some argue that the plural of "octopus" is "octopi" or "octopodes. Does it serve any linguistic purpose? I don't see how, but it won't stop some people for saying it. Look at the use of the subjunctive in "if I were you", which is one of the only occurrences of the subjunctive outside of set-phrases in day-to-day modern English. Is it really necessary? If one was to say "if I was you" would it lose some additional nuance or meaning in practice?

I often see people trying to rationalize some language features (such as arbitrary genders of nouns in many languages) as error correction or some "optimization" but I'm generally unconvinced. Maybe "I went" is just the linguistic equivalent of a platypus, some weird byproduct of a very long evolution with no other intrinsic purpose in the grand scheme of things.

Re: Researchers: It takes 1.5 MB of data to store language information

#45
post #29
post #27

Earlier quoted context omitted.

Aside from that - is there really a way to express neural connections in the brain in terms of bits and bytes?

The problem isn't whether there is "a" way; the problem is there's a multiplicity of ways and we don't know what the best one is. Part of the reason we don't have any idea what the best is is that we can't even construct any of them from a real brain, that is, the "brain scanning" technology the techno-rapture-style Singularity expects to develop does not even remotely exist right now. So we have nothing to experimen…

Brain scanning techniques and simulation software have a long way to go, but there have been significant advances made:

A full fly connectome: https://www.nature.com/articles/s41684-018-0183-8

Architecture of the Mouse Brain Synaptome: https://www.cell.com/neuron/fulltext/S0896-6273(18)30581-6

So we are past just a simple worm simulation now.

Re: Researchers: It takes 1.5 MB of data to store language information

#46

Earlier quoted context omitted.

> There's a good reason to believe evolution would use something close to the most information theoretic efficient implementation. People used to think that about genomes, until we got to know in some detail what's in them.

Explain.

As far as the purpose of DNA is to tell the rest of the cell what proteins to synthesize, over 85% percent of our DNA is never translated to protein [1]. But leave alone our knowledge of DNA, biology rarely works in a manner that's best suited for functional efficiency. For examples, look at how the vagus nerve evolved, or look at vestigial organs. The question for the gene is rarely "is this the most efficient way of doing things?" Rather, it is "is this strategy good enough for me (and my descendants) to pass on my genes?"

[1] https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3205562/

Re: Researchers: It takes 1.5 MB of data to store language information

#47

Earlier quoted context omitted.

> There's a good reason to believe evolution would use something close to the most information theoretic efficient implementation Do you have a citation? I would be very interested to see it. I hate to make analogies between computing and biological systems, but I immediately think of denormalization in databases, error correcting codes, and database replication as examples of optimizations that take you farther from…

Shannon observed this in his work towards information theory. Humans are able to predict next letters in a stream of natural language with a probability approaching optimal. The original paper is here: https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf The summary at wikipedia: https://en.wikipedia.org/wiki/Entropy_(information_theory)#E... Notably, PPM, the best compression algorithm (at the time), was not…

I think we may be talking about different things. You're talking about the language itself. I was talking about the neurological structures that produce the behaviors that are language.

I poked through the optimality literature in biology pretty thoroughly a few years back. There is a great deal of interest, but except in a few cases where a simple piece of natural history is nicely described by an evolutionary game, the choice of criterion to optimize is generally arbitrary. Without an evolutionary mechanism and observations of forces driving that mechanism there is no reason beyond just-so stories to apply any particular criterion.

Re: Researchers: It takes 1.5 MB of data to store language information

#48

Earlier quoted context omitted.

Shannon observed this in his work towards information theory. Humans are able to predict next letters in a stream of natural language with a probability approaching optimal. The original paper is here: https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf The summary at wikipedia: https://en.wikipedia.org/wiki/Entropy_(information_theory)#E... Notably, PPM, the best compression algorithm (at the time), was not…

I think we may be talking about different things. You're talking about the language itself. I was talking about the neurological structures that produce the behaviors that are language. I poked through the optimality literature in biology pretty thoroughly a few years back. There is a great deal of interest, but except in a few cases where a simple piece of natural history is nicely described by an evolutionary game,…

I see the distinction but I am following Chomsky, that they're aspects of the same thing; language is a faculty in a specific, recently evolved neurological structure. That's a bold theory, but it seems to have held up better than approaches where language is an ability learned through a general cognitive intelligence mechanism. Chomsky says in that lecture that it's separate (and compact and recent) and somehow coupled to the general intelligence and performative areas located elsewhere.

I think we're at a point with language and its neural mechanism where by analogy it looks like the phenomenon of flight being closely determined by the exact form and use of a wing. Almost any shape will not do, so there is a form closely (efficiently) following function. Of course there were preadaptations for both which provide very inefficient flight or communication abilities, but something flips and the adaptive utility of the function rapidly sculpts and elaborates the form in that direction.

About optimality, generally agree though I think following thermodynamic efficiency seems a deep inductive step as that allows measure of what is arbitrary or not, e.g. in the work of England that I cited.

Re: Researchers: It takes 1.5 MB of data to store language information

#49
post #44
post #41

Earlier quoted context omitted.

> I took it the other way around: it's easy to memorize "went" because you use it all the time. That still wouldn't explain the why of having it like "went" vs "goed". Sure, it's easy to memorize because we use it all the time, but why have it to memorize it in the first place, versus something like "goed". So, this theory (I tried to convey above) said that it being irregular placed ensured we don't slow down try to…

Your point is interesting but I think you're falling for the same type of fallacies people often have regarding evolution (if you keep going into water for a long time over many generations you'll eventually grow gills!). Natural languages are not designed, they evolve. The irregular nature of the conjugation of "to go" might just be a remnant of some archaic form and nothing else, in the same way some argue that the…

>Natural languages are not designed, they evolve.

But that's part of my point. I don't say irregular verbs were designed to be effective that way, but that they were evolved to be effective.

Hence the link with "language acquisition" being involved -- when regular and irregular verbs developed that wasn't a known theory some "language designer" could consciously follow. Just something innate that could develop because of a evolutionary advantage.

In fact, if someone merely designed, they'd probably go for all regular, rather than regular + irregular, as it's "cleaner".

>I often see people trying to rationalize some language features (such as arbitrary genders of nouns in many languages) as error correction or some "optimization" but I'm generally unconvinced.

Tons of language features are indeed optimizations for different things. Cold climates for example have languages with less vowels (keeping your mouth closed more).

Re: Researchers: It takes 1.5 MB of data to store language information

#50
post #29

Earlier quoted context omitted.

The problem isn't whether there is "a" way; the problem is there's a multiplicity of ways and we don't know what the best one is. Part of the reason we don't have any idea what the best is is that we can't even construct any of them from a real brain, that is, the "brain scanning" technology the techno-rapture-style Singularity expects to develop does not even remotely exist right now. So we have nothing to experimen…

Brain scanning techniques and simulation software have a long way to go, but there have been significant advances made: A full fly connectome: https://www.nature.com/articles/s41684-018-0183-8 Architecture of the Mouse Brain Synaptome: https://www.cell.com/neuron/fulltext/S0896-6273(18)30581-6 So we are past just a simple worm simulation now.

Thank you for the better links. I just meant that as an example.
Post reply on HN