How expensive is a "brute force" approach to decode it? I mean, how about mapping each unknown word by a known word in a known language and improve this mapping until a 'high score' is reached?
That’s a really interesting question — and one I’ve been circling in the back of my head, honestly. I’m not a cryptographer, so I can’t speak to how feasible a brute-force approach is at scale, but the idea of mapping each Voynich “word” to a real word in another language and optimizing for coherence definitely lines up with some of the more experimental approaches people have tried. The challenge (as I understand it…
Show HN: I modeled the Voynich Manuscript with SBERT to test for structure
61–70 of 142 posts
Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure
#62Wasn't it already deciphered, though? https://www.researchgate.net/publication/368991190_The_Voyni...
Most agree that this is not a real solution. Many of the pages translate to nonsense using that scheme, and some of the figures included in the paper don't actually come from the Voynich manuscript in the first place. For more info, see https://www.voynich.ninja/thread-3940-post-53738.html#pid537...
Yet 10 years later I still hear that the consensus is that there's no agreeable translation. So, what, all this mandaic-gypsies was nothing? And all coincidences were… coincidences?
Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure
#63Also there might be some characters that are in there just to confuse. For example that bizarre capital "P"-like thing that has multiple variations seems to appear sometimes far too often to represent real language, so it might be just an obfuscator that's removed prior to decryption. There may be other characters that are abnormally "frequent" and they're maybe also unused dummy characters. But the "too many Ps" problem is also consistent with just pure fiction too, I realize.
Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure
#64Earlier quoted context omitted.
Statistical analyses such as this one consistently find patterns that are consistent with a proper language and would be unlikely to have emerged from someone who was just putting gibberish on the page. To get the kinds of patterns these turn up someone would have had to go a large part of the way towards building a full constructed language, which is interesting in its own right.
Even before we consider the cipher, there's a huge difference between a constructed language and a stochastic process to generate language-like text.
Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure
#65How expensive is a "brute force" approach to decode it? I mean, how about mapping each unknown word by a known word in a known language and improve this mapping until a 'high score' is reached?
That’s a really interesting question — and one I’ve been circling in the back of my head, honestly. I’m not a cryptographer, so I can’t speak to how feasible a brute-force approach is at scale, but the idea of mapping each Voynich “word” to a real word in another language and optimizing for coherence definitely lines up with some of the more experimental approaches people have tried. The challenge (as I understand it…
Maybe a version of scripture that had been "rejected" by some King, and was illegal to reproduce? Take the best radiocarbon dating, figure out who was King back then, and if they 'sanctioned' any biblical translations, and then go to the version of the bible before that translation, and this will be what was perhaps illegal and needed to be encrypted. That's just one plausible story. Who knows, we might find out the phrase "young girl" was simplified to "virgin", and that would potentially be a big secret.
Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure
#66Earlier quoted context omitted.
When I get nice separation with PCA, I personally tend to eschew UMAP, since the relative distance of all the points to one another is easier to interpret. I avoid t-SNE at all costs, because distance in those plots are pretty much meaningless. (Before I get yelled out, this isn't prescriptive, it's a personal preference.)
PCA having nice separation is extremely uncommon unless your data is unusually clean or has obvious patterns. Even for the comically-easy MNIST dataset, the PCA representation doesn't separate nicely: https://github.com/lmcinnes/umap_paper_notebooks/blob/master...
I'd add that just because you can achieve separability from a method, the resulting visualization may not be super informative. The distance between clusters that appear in t-SNE projected space often have nothing to do with their distance in latent space, for example. So while you get nice separate clusters, it comes at the cost of the projected space greatly distorting/hiding the relationship between points across clusters.
Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure
#67I strongly believe the manuscript is undecipherable in the sense thats it's all gibberish. I can't prove it, but at this point I think it's more likely than not to be hoax.
Statistical analyses such as this one consistently find patterns that are consistent with a proper language and would be unlikely to have emerged from someone who was just putting gibberish on the page. To get the kinds of patterns these turn up someone would have had to go a large part of the way towards building a full constructed language, which is interesting in its own right.
Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure
#68TIL about the Voynich manuscript. Fascinating. Thank you.
Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure
#69Earlier quoted context omitted.
Statistical analyses such as this one consistently find patterns that are consistent with a proper language and would be unlikely to have emerged from someone who was just putting gibberish on the page. To get the kinds of patterns these turn up someone would have had to go a large part of the way towards building a full constructed language, which is interesting in its own right.
If you're going to make a hoax for fun or for profit, wouldn't it be the best first step to make it seem legitimate, by coming up with a fake language? Klingon is fake, but has standard conventions. This isn't really a difficult proposition compared to all of the illustrations and what-not, I would think.
Re: Show HN: I modeled the Voynich Manuscript with SBERT to test for structure
#70what I'd expect from a handwritten book like that, if it is just a gibberish, and not a cypher of any sorts - the style, calligraphy, the words used, even letters themselves should evolve from page 1 to the last page. Pages could be reordered of course, but it still should be noticeable. Unless author hadn't written tens of books exactly like that before, which didn't survive, of course. I don't think it's a very nov…
A lot of work's been done here. There are believed to have been 2 scribes (see Prescott Currier), although Lisa Fagin Davis posits 5. Here's a discussion of an experiment working off of Fagin Davis' position: https://www.voynich.ninja/thread-3783.html