Live data from Hacker News

Show HN: I modeled the Voynich Manuscript with SBERT to test for structure

github.com

21–30 of 142 posts

Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure

#21
The best work on Voynich has been done by Emma Smith, Coons and Patrick Feaster, about loops and QOKEDAR and CHOLDAIIN cycles. Here's a good presentation: https://www.youtube.com/watch?v=SCWJzTX6y9M Zattera and Roe have also done good work on the "slot alphabet". That so many are making progression in the same direction is quite encouraging!

https://www.voynich.ninja/thread-4327-post-60796.html#pid607... is the main forum discussing precisely this. I quite liked this explanation of the apparent structure: https://www.voynich.ninja/thread-4286.html

> RU SSUK UKIA UK SSIAKRAINE IARAIN RA AINE RUK UKRU KRIA UKUSSIA IARUK RUSSUK RUSSAINE RUAINERU RUKIA

That is, there may be 2 "word types" with different statistical properties (as Feaster's video above describes)(perhaps e.g. 2 different Cyphers used "randomly" next to each other). Figuring out how to imitate the MS' statistical properties would let us determine cypher system and make steps towards determining its language etc. so most credible work's gone in this direction over the last 10+ years.

This site is a great introduction/deep dive: https://www.voynich.nu/

Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure

#22
post #21

The best work on Voynich has been done by Emma Smith, Coons and Patrick Feaster, about loops and QOKEDAR and CHOLDAIIN cycles. Here's a good presentation: https://www.youtube.com/watch?v=SCWJzTX6y9M Zattera and Roe have also done good work on the "slot alphabet". That so many are making progression in the same direction is quite encouraging! https://www.voynich.ninja/thread-4327-post-60796.html#pid607... is the main…

I’m definitely not a Voynich expert or linguist — I stumbled into this more or less by accident and thought it would make for a fun NLP learning project. Really appreciate you pointing to those names and that forum — I wasn’t aware of the deeper work on QOKEDAR/CHOLDAIIN cycles or the slot alphabet stuff. It’s encouraging to hear that the kind of structure I modeled seems to resonate with where serious research is heading.

Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure

#23
post #15

I strongly believe the manuscript is undecipherable in the sense thats it's all gibberish. I can't prove it, but at this point I think it's more likely than not to be hoax.

Statistical analyses such as this one consistently find patterns that are consistent with a proper language and would be unlikely to have emerged from someone who was just putting gibberish on the page. To get the kinds of patterns these turn up someone would have had to go a large part of the way towards building a full constructed language, which is interesting in its own right.

Could still be gibberish.

Shud less kee chicken souls do be gooby good? Mus hess to my rooby roo!

Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure

#24
post #7

UMAP or TSNE would be nice, even if PCA already shows nice separation. Reference mapping each cluster to all the others would be a nice way to indicate that there's no variability left in your analysis

Do you have examples of how this reference mapping is performed? I'm interested in this for embeddings in a different modality, but don't have as much experience on the NLP side of things

Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure

#25

I’ve found this to be one of the most interesting hypotheses: http://voynichproject.org/ The author made an assumption that Voynichese is a Germanic language, and it looks like he was able to make some progress with it. I’ve also come across accounts that it might be an Uralic or Finno-Ugric language. I think your approach is great, and I wonder if tweaking it for specific language families could go even further.

This thread discusses the many purported "solutions": https://www.voynich.ninja/thread-4341.html While Bernholz' site is nice, Child's work doesn't shed much light on actually deciphering the MS.

Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure

#26
post #25

I’ve found this to be one of the most interesting hypotheses: http://voynichproject.org/ The author made an assumption that Voynichese is a Germanic language, and it looks like he was able to make some progress with it. I’ve also come across accounts that it might be an Uralic or Finno-Ugric language. I think your approach is great, and I wonder if tweaking it for specific language families could go even further.

This thread discusses the many purported "solutions": https://www.voynich.ninja/thread-4341.html While Bernholz' site is nice, Child's work doesn't shed much light on actually deciphering the MS.

Thanks for this! I had come across Child’s hypothesis after doing a search related to Old Prussian and Slavic languages, so I don’t have much context for this solution, and this is helpful to see.

Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure

#27
post #7

UMAP or TSNE would be nice, even if PCA already shows nice separation. Reference mapping each cluster to all the others would be a nice way to indicate that there's no variability left in your analysis

When I get nice separation with PCA, I personally tend to eschew UMAP, since the relative distance of all the points to one another is easier to interpret. I avoid t-SNE at all costs, because distance in those plots are pretty much meaningless. (Before I get yelled out, this isn't prescriptive, it's a personal preference.)

We are of a like mind.

Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure

#28
post #21

The best work on Voynich has been done by Emma Smith, Coons and Patrick Feaster, about loops and QOKEDAR and CHOLDAIIN cycles. Here's a good presentation: https://www.youtube.com/watch?v=SCWJzTX6y9M Zattera and Roe have also done good work on the "slot alphabet". That so many are making progression in the same direction is quite encouraging! https://www.voynich.ninja/thread-4327-post-60796.html#pid607... is the main…

Ock ohem octei wies barsoom?

Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure

#29
post #15

I strongly believe the manuscript is undecipherable in the sense thats it's all gibberish. I can't prove it, but at this point I think it's more likely than not to be hoax.

Statistical analyses such as this one consistently find patterns that are consistent with a proper language and would be unlikely to have emerged from someone who was just putting gibberish on the page. To get the kinds of patterns these turn up someone would have had to go a large part of the way towards building a full constructed language, which is interesting in its own right.

> consistent with a proper language

There's certainly a system to the madness, but it exhibits rather different statistical properties from "proper" languages. Look at section 2.4: https://www.voynich.nu/a2_char.html At the moment, any apparently linguistic patterns are happenstance; the cypher fundamentally obscures its actual distribution (if a "proper" language.)

Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure

#30
post #15

I strongly believe the manuscript is undecipherable in the sense thats it's all gibberish. I can't prove it, but at this point I think it's more likely than not to be hoax.

Statistical analyses such as this one consistently find patterns that are consistent with a proper language and would be unlikely to have emerged from someone who was just putting gibberish on the page. To get the kinds of patterns these turn up someone would have had to go a large part of the way towards building a full constructed language, which is interesting in its own right.

> would be unlikely to have emerged from someone who was just putting gibberish on the page

People often assert this, but I'm unsure of any evidence. If I wrote a manuscript in a pretend language, I would expect it to end up with language-like patterns, some automatically and some intentionally.

Humans aren't random number generators, and they aren't stupid. Therefore, the implicit claim that a human could not create a manuscript containing gibberish that exhibits many language-like patterns seems unlikely to be true.

So we have two options:

1. This is either a real language or an encoded real language that we've never seen before and can't decrypt, even after many years of attempts

2. Or it is gibberish that exhibits features of a real language

I can't help but feel that option 2 is now the more likely choice.

Post reply on HN