Live data from Hacker News

Federated finetuning of Whisper on Raspberry Pi 5

flower.dev

1–10 of 21 posts

Re: Federated finetuning of Whisper on Raspberry Pi 5

#4
How would this actually work in practice? Do I ask the user to utter specific words then train on that? How is it different from the traditional speech recognition that I need to 'train' to work better on my voice?

The Holy Grail would be to train the model while using it, without any friction. I don't think these methods support that though.

Re: Federated finetuning of Whisper on Raspberry Pi 5

#5

I don't think the article mentions it, how well does the rpi 4 and 5 do for inference with whisper especially v3?

I’m also interested in peoples’ experience. I’d expect decent performance: Whisper 3 has many model sizes, down to 35Mb, iirc. Training, and especially inference, should be doable on a Pi5.

Re: Federated finetuning of Whisper on Raspberry Pi 5

#7

How would this actually work in practice? Do I ask the user to utter specific words then train on that? How is it different from the traditional speech recognition that I need to 'train' to work better on my voice? The Holy Grail would be to train the model while using it, without any friction. I don't think these methods support that though.

One of the Flower maintainers here. The code example is primarily meant as a demonstrator to show that it's possible to fine-tune these models in a federated way on devices as small as a Raspberry Pi 5.

The bigger takeaway is that we're close to being able to train/fine-tune models with much better performance by accessing vastly more data on the edge, in a federated way.

Re: Federated finetuning of Whisper on Raspberry Pi 5

#8
post #3

I’m guessing this will also help with thick accents?

yeah. with FL it should be possible to make sense out of all data that is distributed across devices without ever having to move it to a central location (i.e. collect it). In the case of speech data, users participating in a federated setting would likely come from different backgrounds, which could be reflected in their accent or use of language.

Re: Federated finetuning of Whisper on Raspberry Pi 5

#9

I don't think the article mentions it, how well does the rpi 4 and 5 do for inference with whisper especially v3?

v3 only comes in one flavor: large.

I don’t think you’re going to have a good time running the large model on a Pi of any kind.

The large models are 32x slower than the tiny models, roughly.[0]

I just tested, and whisper.cpp on my Pi 4 can transcribe the 30-second a13.wav sample (“make samples” to fetch it) in 18.5 seconds.

You can do the math… 32x = 10 minutes transcribe 30 seconds of audio with the large model. Not a good time for most people.

The Pi 5 could be 2x to 3x faster.

[0]: https://github.com/openai/whisper/blob/main/README.md#availa...

Re: Federated finetuning of Whisper on Raspberry Pi 5

#10
post #9

I don't think the article mentions it, how well does the rpi 4 and 5 do for inference with whisper especially v3?

v3 only comes in one flavor: large. I don’t think you’re going to have a good time running the large model on a Pi of any kind. The large models are 32x slower than the tiny models, roughly.[0] I just tested, and whisper.cpp on my Pi 4 can transcribe the 30-second a13.wav sample (“make samples” to fetch it) in 18.5 seconds. You can do the math… 32x = 10 minutes transcribe 30 seconds of audio with the large model. Not…

I can confirm that we're seeing 2x to 3x faster (RPi 4 vs RPi 5) in some of our early tests
Post reply on HN