Live data from Hacker News

Researchers: It takes 1.5 MB of data to store language information

medicalxpress.com

51–59 of 59 posts

Re: Researchers: It takes 1.5 MB of data to store language information

#51
post #46

Earlier quoted context omitted.

Explain.

As far as the purpose of DNA is to tell the rest of the cell what proteins to synthesize, over 85% percent of our DNA is never translated to protein [1]. But leave alone our knowledge of DNA, biology rarely works in a manner that's best suited for functional efficiency. For examples, look at how the vagus nerve evolved, or look at vestigial organs. The question for the gene is rarely "is this the most efficient way o…

That's a pretty cursory understanding of how DNA works. The "Junk DNA" hypothesis has long been invalidated.

The non-coding genome is a large part of what causes us to look different from a mouse or a fruit-fly. Gene regulation (what gene to transcribe, and when) is largely controlled by non-coding DNA: https://en.wikipedia.org/wiki/Regulation_of_gene_expression

There are energy costs associated with useless information being stored and transcribed in the genome... so it is pretty reasonable to believe that information theory does play some (however minor) role in how data is stored.

Re: Researchers: It takes 1.5 MB of data to store language information

#52
post #46

Earlier quoted context omitted.

As far as the purpose of DNA is to tell the rest of the cell what proteins to synthesize, over 85% percent of our DNA is never translated to protein [1]. But leave alone our knowledge of DNA, biology rarely works in a manner that's best suited for functional efficiency. For examples, look at how the vagus nerve evolved, or look at vestigial organs. The question for the gene is rarely "is this the most efficient way o…

That's a pretty cursory understanding of how DNA works. The "Junk DNA" hypothesis has long been invalidated. The non-coding genome is a large part of what causes us to look different from a mouse or a fruit-fly. Gene regulation (what gene to transcribe, and when) is largely controlled by non-coding DNA: https://en.wikipedia.org/wiki/Regulation_of_gene_expression There are energy costs associated with useless informat…

We are drifting away from the original issue here. 'However minor' is a long way from near-maximal efficiency, and in the context of the original claim, 'efficient' means compact encoding.

The fact that the DNA involved in gene regulation is mostly non-coding does not mean that most non-coding DNA has a regulatory purpose. If it did, changes in non-coding DNA would be more important, in aggregate, than changes in coding DNA are, with respect to evolution and disease, on account of its prevalence.

Re: Researchers: It takes 1.5 MB of data to store language information

#53
post #46

Earlier quoted context omitted.

As far as the purpose of DNA is to tell the rest of the cell what proteins to synthesize, over 85% percent of our DNA is never translated to protein [1]. But leave alone our knowledge of DNA, biology rarely works in a manner that's best suited for functional efficiency. For examples, look at how the vagus nerve evolved, or look at vestigial organs. The question for the gene is rarely "is this the most efficient way o…

That's a pretty cursory understanding of how DNA works. The "Junk DNA" hypothesis has long been invalidated. The non-coding genome is a large part of what causes us to look different from a mouse or a fruit-fly. Gene regulation (what gene to transcribe, and when) is largely controlled by non-coding DNA: https://en.wikipedia.org/wiki/Regulation_of_gene_expression There are energy costs associated with useless informat…

Neither I nor the paper I cited use the phrase "Junk DNA".

> The non-coding genome is a large part of what causes us to look different from a mouse or a fruit-fly.

Now that's just fantasy. There have been a lot of changes since our common ancestors with mice and fruit flies. We are all animals and obviously would share genes that perform analogous functions -- like common genes for controlling wing shapes and organ shapes. But would you like to point me to evidence in the human DNA that codes for fruit fly wings? And what role do you think having virus DNA in our genome plays in our development?

DNA encodes information. Information theory does play a role. But as mannykannot pointed out, the topic of the conversation is something else.

Re: Researchers: It takes 1.5 MB of data to store language information

#54

Earlier quoted context omitted.

That's a pretty cursory understanding of how DNA works. The "Junk DNA" hypothesis has long been invalidated. The non-coding genome is a large part of what causes us to look different from a mouse or a fruit-fly. Gene regulation (what gene to transcribe, and when) is largely controlled by non-coding DNA: https://en.wikipedia.org/wiki/Regulation_of_gene_expression There are energy costs associated with useless informat…

We are drifting away from the original issue here. 'However minor' is a long way from near-maximal efficiency, and in the context of the original claim, 'efficient' means compact encoding. The fact that the DNA involved in gene regulation is mostly non-coding does not mean that most non-coding DNA has a regulatory purpose. If it did, changes in non-coding DNA would be more important, in aggregate, than changes in cod…

Junk DNA is a phrase biologists use for describing the long held (and recently disproven) idea that non-coding DNA ("dna not coding for proteins") is unimportant relative to coding regions (genes).

Most non-coding DNA does have a regulatory purpose. Things like copy number variations (where information is "duplicated") between promoters and enhancers can can change the rate at which particular genes get transcribed. This is my point: that viewing the genome as "information" in the simple sense of each ATCG nucleotide corresponding to bits is absurd. The genome is a geometric AND chemical entity and encodes information using both paradigms.

My point is not to say that you are wrong, but rather that you do not know. It is FAR from well established that the genome is inefficient at storing information.

Re: Researchers: It takes 1.5 MB of data to store language information

#55
post #19

I believe it if they provide a program that can speak English in under 1.5MB total size.

What does it mean to "speak English"? Does a program that plays a recording of the word "I" count as a program that speaks English?

I think you know the answer: think about how you find out whether a human speaks English or not.

Re: Researchers: It takes 1.5 MB of data to store language information

#56
post #4

How would it be possible that 40,000 words translates to 400,000 bits?! Am I missing something here?

A quick test using a non-lossy compressor with no understanding of phonemes or human language and grammatical context at all on a dictionary of 370000 English words resulted in 24 bits per word here. It wouldn't surprise me if our abilities to roughly contextualize language in terms of the language we already understand gives us a serious advantage here.

Now a few questions: Can you hear a word you've never written (in a language that you're familiar with) and intuitively spell it right the first time? Can you read a word that you don't immediately understand and figure out its meaning from the context in which it is used? Can you accurately complete half of a sentence?

That a lot of people can do these things suggests to me that we all sit on something superficially similar to an efficient lossy compressor in our brain.

Re: Researchers: It takes 1.5 MB of data to store language information

#57

Earlier quoted context omitted.

We are drifting away from the original issue here. 'However minor' is a long way from near-maximal efficiency, and in the context of the original claim, 'efficient' means compact encoding. The fact that the DNA involved in gene regulation is mostly non-coding does not mean that most non-coding DNA has a regulatory purpose. If it did, changes in non-coding DNA would be more important, in aggregate, than changes in cod…

Junk DNA is a phrase biologists use for describing the long held (and recently disproven) idea that non-coding DNA ("dna not coding for proteins") is unimportant relative to coding regions (genes). Most non-coding DNA does have a regulatory purpose. Things like copy number variations (where information is "duplicated") between promoters and enhancers can can change the rate at which particular genes get transcribed.…

I don't know, you don't know, but only one of us is using speculative arguments (and, in at least one earlier case, fantasy, as lake99 pointed out above) to support her opinion.

From an information-theoretic point of view (which is where we started) even if all non-coding DNA had some purpose, it would not be a sufficient basis for assuming that it is close to the most information-theoretically efficient implementation. On the other hand, the fact that whole genomes can be handily compressed, with straightforward Huffman encoding techniques, settles the infomation-theoretic efficiency question, regardless of whether your speculation about the usefulness of all non-coding DNA is correct.

Re: Researchers: It takes 1.5 MB of data to store language information

#58

Earlier quoted context omitted.

I think we may be talking about different things. You're talking about the language itself. I was talking about the neurological structures that produce the behaviors that are language. I poked through the optimality literature in biology pretty thoroughly a few years back. There is a great deal of interest, but except in a few cases where a simple piece of natural history is nicely described by an evolutionary game,…

I see the distinction but I am following Chomsky, that they're aspects of the same thing; language is a faculty in a specific, recently evolved neurological structure. That's a bold theory, but it seems to have held up better than approaches where language is an ability learned through a general cognitive intelligence mechanism. Chomsky says in that lecture that it's separate (and compact and recent) and somehow coup…

> I am following Chomsky, that they're aspects of the same thing; language is a faculty in a specific, recently evolved neurological structure.

That is a bold hypothesis, certainly, but not one that you could use as a basis for argument at this point in time. I also deeply doubt it when viewed in the context of Wittgenstein's language games. A human and a sheep dog can very effectively engage in a language game, even subtle ones such as naming. Nor is it pure behavioral conditioning, as what makes various dog breeds better at particular things is selection for incidence of particular behaviors.

Then you go to a gorilla or a parrot, which can engage in quite abstract language games. Language games with parrots always make you aware of Wittgenstein's lion, though.

For Chomsky's hypothesis to be true, you would need to have a particular inflection point where language games become language, and for that inflection point to be reified as the evolution of a specific neural structure instead of exaptation of lots of aspects in the brain.

It feels too much like Watson and Crick's nonsense "one gene = one protein," which could have been discarded with a modest amount of thought before it was allowed to damage biology for decades.

> I think following thermodynamic efficiency seems a deep inductive step as that allows measure of what is arbitrary or not

Until you look at, say, bower birds and realize that conspicuous waste is a signalling mechanism for fitness. There you have an evolutionary game that rewards thermodynamic inefficiency.

Re: Researchers: It takes 1.5 MB of data to store language information

#59
post #36

Earlier quoted context omitted.

English cognate to "atreve" is "attribute". ~ов is an iterative suffix - like ~le or ~er in gamble and chatter. It's useful to know, because you can rationalise why it is always dropped in the present tense - you can't be iterative at the moment, unless you're an Englishman)

>English cognate to "atreve" is "attribute". I didn't know that, but you'll grant me that it's not a very useful cognate (either in spelling or in meaning). >~ов is an iterative suffix - like ~le or ~er in gamble and chatter. It's useful to know, because you can rationalise why it is always dropped in the present tense - you can't be iterative at the moment, unless you're an Englishman) Very interesting, thanks.

If you're serious about russian, feel free to ping me.
Post reply on HN