Live data from Hacker News

Vāgdhenu: A Sanskrit Chanting TTS System

prathosh.in

21–30 of 65 posts

Re: Vāgdhenu: A Sanskrit Chanting TTS System

#21
post #13

The TTS quality seems to be poor. I just entered a single word "jaganmaatham" and it failed to pronounce it. What's the point anyway? If you really love your traditions, slokas and chanting, please keep it as natural as possible, instead of plasticizing it to the core. This is third level of mechanizing the sacred rituals. First - we lost the live chanting due to recorded recitals with human voice. Then we lost the e…

Yours is a very shortsighted and myopic view on the importance of using modern technology in order to maintain/enhance/propagate ancient cultural/linguistic knowledge which is losing its appeal in the "common populace".

This is another tool to help correct that balance.

The author seems a pretty knowledgeable person - https://prathosh.in/

His talk "The Non-negotiability of Shastra-s in Modern Science" - https://www.youtube.com/watch?v=i5Tp4gdQiTM

Re: Vāgdhenu: A Sanskrit Chanting TTS System

#22
post #18
post #13

The TTS quality seems to be poor. I just entered a single word "jaganmaatham" and it failed to pronounce it. What's the point anyway? If you really love your traditions, slokas and chanting, please keep it as natural as possible, instead of plasticizing it to the core. This is third level of mechanizing the sacred rituals. First - we lost the live chanting due to recorded recitals with human voice. Then we lost the e…

Seems like its point is to help people practice themselves. > Vāgbodhinī A Sanskrit chant tutor. Paste any śloka (or prose) in any script, hear its metre-aware reference chant rendered by Vāgdhenu, then chant along — a Sanskrit speech model scores every syllable and shows you what to fix. For classical (laukika) Sanskrit.

That's how they will present it. But the direction is clear. It can replace a human chanting professional. If it can, it will. The future is not far off where an AI bot will interact with the users, asking their family name etc and weave it into the chanting.

Re: Vāgdhenu: A Sanskrit Chanting TTS System

#23
post #13

The TTS quality seems to be poor. I just entered a single word "jaganmaatham" and it failed to pronounce it. What's the point anyway? If you really love your traditions, slokas and chanting, please keep it as natural as possible, instead of plasticizing it to the core. This is third level of mechanizing the sacred rituals. First - we lost the live chanting due to recorded recitals with human voice. Then we lost the e…

Yours is a very shortsighted and myopic view on the importance of using modern technology in order to maintain/enhance/propagate ancient cultural/linguistic knowledge which is losing its appeal in the "common populace". This is another tool to help correct that balance. The author seems a pretty knowledgeable person - https://prathosh.in/ His talk "The Non-negotiability of Shastra-s in Modern Science" - https://www.y…

You don't need to know about the person, to understand what is done here. This is no invention or profound deed. Anyone using the open-weights TTS models can do this. Talk about the overall goal.

Re: Vāgdhenu: A Sanskrit Chanting TTS System

#25
post #20
post #13

The TTS quality seems to be poor. I just entered a single word "jaganmaatham" and it failed to pronounce it. What's the point anyway? If you really love your traditions, slokas and chanting, please keep it as natural as possible, instead of plasticizing it to the core. This is third level of mechanizing the sacred rituals. First - we lost the live chanting due to recorded recitals with human voice. Then we lost the e…

I have tried the word you mentioned along with few other words together and there were no issues. It doesn't have to be a replacement. Anyone who has ever been in the physical presence of an expert reciting sanskrit verses knows that this can never replace that experience. It will only make the language and the chants more accessible.

> I have tried the word you mentioned along with few other words together and there were no issues.

I have typed the word in English literals, and it made a horrible pronunciation of the word.

> It will only make the language and the chants more accessible.

No, it will turn sacred chanting into an AI slop. Entirely due to desire of people wishing to make some quick bucks or cheap reputation.

Re: Vāgdhenu: A Sanskrit Chanting TTS System

#26
post #14
post #8

It was based on an existing TTS, IndicT5. I wonder how different is “Sanskrit Chanting” to languages it could already do, like Hindi. Is it largely the glyph—to-phoneme that needs relearning? Or pitch control? Or more?

The biggest problem with a lot of northern languages is the awful schwa deletion. Even the Marathi of Maharashtra is subject to it. Strangely, the Marathi of Thanjavur has far, far less of it from what little I have heard of it. Sanskrit does not do schwa deletion. So engines must map the phonemes to Kannada/Telugu (which support the full complement of Sanskrit sounds) if they want to get somewhere. A lot of this can…

Luckily a lot of research has gone into Vedic recitation from an academic perspective if someone would take up the mantle.

Re: Vāgdhenu: A Sanskrit Chanting TTS System

#28

Why does the visarga "echo" at the end of the verse? A mixed IAST - HK transliteration on the front page like "Bhāgavata-VāNi" does not inspire confidence either. Is it vibed?

Even if it was vibe-coded, does that matter? The result is solid. If it’s not to your taste, wait for the next release.

Re: Vāgdhenu: A Sanskrit Chanting TTS System

#29
post #23

Earlier quoted context omitted.

Yours is a very shortsighted and myopic view on the importance of using modern technology in order to maintain/enhance/propagate ancient cultural/linguistic knowledge which is losing its appeal in the "common populace". This is another tool to help correct that balance. The author seems a pretty knowledgeable person - https://prathosh.in/ His talk "The Non-negotiability of Shastra-s in Modern Science" - https://www.y…

You don't need to know about the person, to understand what is done here. This is no invention or profound deed. Anyone using the open-weights TTS models can do this. Talk about the overall goal.

You do need to know about a person to understand why he did a project that he did. What were his motivations and goals? That is the point here which can be inferred from his bio/talks.

His goals need not be what you want them to be; his project; his rules.

Nobody is claiming this is something earth-shattering; just that it is an interesting project on a difficult subject matter. The author has linked to the paper for the project on the site itself which you can read for your edification.

From the introduction;

Classical Sanskrit recitation, parāyaṇa, is a chanted rather than a read register. A faithful synthesizer must hold long vowels, sustain a terminal visarga, articulate retroflex and aspirated consonants, render dense consonant conjuncts cleanly, and respect the metrical structure of the verse.

None of these is well served by general-purpose text-to-speech, and there is essentially no chant-domain training data available off the shelf. The problem is therefore doubly hard: it is low-resource, and the target prosody is a specialized melodic contour rather than ordinary read speech.

This report describes a system, Vāgdhenu, that solves the practical version of this problem well enough to ship two large deployments, and it documents the design decisions, the dead ends, and the one negative result that turned out to be the most useful finding. We do not claim a new model. We claim an honest account of what it takes to build a faithful Sanskrit chant pipeline on top of current open backbones, what works, and what is architecturally out of reach.

Our framing is that of an experience report. The evidence we offer is the comparative lineage across architecture families, a reproducible production system, two shipped artifacts at real scale, and a public release of code, weights, data, and a live demonstration. Formal listening studies are limited to expert evaluation, which we state plainly and treat as a limitation rather than a result.

As a practical example; if you have ever heard even the common Gayathri Mantra being chanted by a Tamil Thanjavur Brahmin priest vs. a UP Allahabad Brahmin priest you would know the wide difference in utterances. Having something like this gives you a reference (albeit with a South Indian Kannada bent) to compare against.

Re: Vāgdhenu: A Sanskrit Chanting TTS System

#30
People seem to be making all sorts of assumptions/conjectures about the project and confusing themselves without reading the paper explaining the project (linked to in the site itself);

From the introduction;

Classical Sanskrit recitation, parāyaṇa, is a chanted rather than a read register. A faithful synthesizer must hold long vowels, sustain a terminal visarga, articulate retroflex and aspirated consonants, render dense consonant conjuncts cleanly, and respect the metrical structure of the verse.

None of these is well served by general-purpose text-to-speech, and there is essentially no chant-domain training data available off the shelf. The problem is therefore doubly hard: it is low-resource, and the target prosody is a specialized melodic contour rather than ordinary read speech.

This report describes a system, Vāgdhenu, that solves the practical version of this problem well enough to ship two large deployments, and it documents the design decisions, the dead ends, and the one negative result that turned out to be the most useful finding. We do not claim a new model. We claim an honest account of what it takes to build a faithful Sanskrit chant pipeline on top of current open backbones, what works, and what is architecturally out of reach.

Our framing is that of an experience report. The evidence we offer is the comparative lineage across architecture families, a reproducible production system, two shipped artifacts at real scale, and a public release of code, weights, data, and a live demonstration. Formal listening studies are limited to expert evaluation, which we state plainly and treat as a limitation rather than a result.

Post reply on HN