Live data from Hacker News

Build your own Siri locally and on-device

thehyperplane.substack.com

41–50 of 50 posts

Re: Build your own Siri locally and on-device

#41
post #2

So build your own crappy agent-assistant? In earnest though, I'm certain we'll see a community replacement of Siri by end-of-year if the iPhone permissions model allows it or there's some workaround. IDK what the limitations are here but I'm eagerly awaiting the community to step in where Siri has failed.

A better one* Siri can't even tell me the weather right, this assistant is an elevated version that works first, then performs great second

Re: Build your own Siri locally and on-device

#43

I love the idea and I would like to build something like this. But the few attempts i have made using whisper locally has so far been underwhelming. Has anyone gotten results with small whisper models that are good enough for a use case like this? Maybe I've just had a bad microphone.

Which model do you use? I use large usually, on a GPU. It's fast and works really well. Be aware though that it can only recognise one language at a time. It will autodetect if you don't specify one.

Of course the smaller models don't work nearly as well and they are often restricted to English. Large works great for me though it does require GPU hardware to be responsive enough, even with faster-whisper or insanely-fast-whisper.

Re: Build your own Siri locally and on-device

#44

Does Apple even allow you to replace Siri with another assistant? For the longest time on android, all non-Google assistants were crippled by not being able to listen in the background or use the assistant hardkey, gestures, or shortcuts. I'm not sure if the Google assistant still has privileges others don't, but I wouldn't be surprised in the least.

Part of the problem is the wake word “hey siri” is actually handed by a separate coprocessor (AOP) with the model compiled into the firmware. While anything is technically possible, it isn’t as simple as just letting the google app run in the background since the AP is asleep when any of these gesture happen. You could probably setup the action button on the side to open an assistant, but that’s going to be a less pl…

There's open solutions for that like openwakeword and microwakeword (the latter can even run on an esp32!)

The training is a lot of work though and requires a lot of material. For Home Assistant's voice preview model they had tens of thousands of volunteers record the "okay nabu" wakeword and even still it doesn't work quite as well as hey siri on Apple devices.

Re: Build your own Siri locally and on-device

#45

Earlier quoted context omitted.

Crappy? Dude, Siri at one point couldn't even tell you what today's date was. The bar is on the ground.

I do not know if it is because I have been trained to make simple requests, but there are only a half dozen things I would verbally ask a robot. - time of day - calendar date - weather - set a timer - simple math calculation That’s 90% of the functionality right there.

For me I ask a lot of things like "How do I say in Spanish". It's better than a google translate because it's not quite as literal, it will translate to proper colloquialisms if necessary.

Re: Build your own Siri locally and on-device

#46
post #28

Earlier quoted context omitted.

+1 this. Whisper works insanely well. I've been using the medium model as it has yet to mis transcribe anything noticeable, and it's very lightweight. I even converted it to a coreML model so it runs accelerated on apple silicon. It doesn't run *that* much faster than before.. but it ran really fast to begin with. For anyone tinkering, ive had much success with whisper.cpp.

What was the process of converting it like? I assume you then had to write all of the inference code as well?

not the gp but found this https://github.com/ggml-org/whisper.cpp/blob/master/models/c...

Re: Build your own Siri locally and on-device

#47

Does Apple even allow you to replace Siri with another assistant? For the longest time on android, all non-Google assistants were crippled by not being able to listen in the background or use the assistant hardkey, gestures, or shortcuts. I'm not sure if the Google assistant still has privileges others don't, but I wouldn't be surprised in the least.

Part of the problem is the wake word “hey siri” is actually handed by a separate coprocessor (AOP) with the model compiled into the firmware. While anything is technically possible, it isn’t as simple as just letting the google app run in the background since the AP is asleep when any of these gesture happen. You could probably setup the action button on the side to open an assistant, but that’s going to be a less pl…

You can now setup Vocal Shortcuts[1] which can be used to run any shortcut or action with almost any trigger word and without saying "Siri". However, I'm not certain if it can wake the device from sleep or not.

[1] https://support.apple.com/en-in/guide/iphone/iph7f242ea2c/io...

Re: Build your own Siri locally and on-device

#48

Earlier quoted context omitted.

This summary-like style — with heavy formatting and every (!) paragraph as a bulleted list — drives me nuts tbqh. Especially in lengthy texts, it just looks... noisy, bland, and sometimes confusing.

What's the format you would prefer? We're not using ChatGPT to write and we've experimented with this format. The other articles may have a better format?

> We're not using ChatGPT to write

We can go into semantics, but that listicle has all the hallmarks of LLM-produced stuff, down the the "Generated Image"-tagged image.

Re: Build your own Siri locally and on-device

#49

I love the idea and I would like to build something like this. But the few attempts i have made using whisper locally has so far been underwhelming. Has anyone gotten results with small whisper models that are good enough for a use case like this? Maybe I've just had a bad microphone.

> Maybe I've just had a bad microphone. Yeah, I would definitely double-check your setup. At work we use Whisper to live-transcribe-and-translate all-hands meetings and it works exceptionally well.

I'd agree with your experience. I simply sit my phone (~200 dollar motorola, cheap phone) in centre of room, split voice file into chunks using voice prints/ID's I get from a voice embedding model I trained, then feed labelled chunks through whisper, and get a nice transcript of everything said. I combine that with my handwritten notes (as image, get a VLM to transcribe) and the agenda, and I get out really nice meeting minutes as a LaTex document. Works a charm and has turned an hour or two of work per meeting into maybe 30 minutes (proofing what was written).

Re: Build your own Siri locally and on-device

#50

Earlier quoted context omitted.

What's the format you would prefer? We're not using ChatGPT to write and we've experimented with this format. The other articles may have a better format?

> We're not using ChatGPT to write We can go into semantics, but that listicle has all the hallmarks of LLM-produced stuff, down the the "Generated Image"-tagged image.

The image yes, it was generated and that's pretty clear.

We used to rely more on ChatGPT but that did not work out (figures!) so now it's a lot of authentic human speech.

So I'm curious when you say it has the hallmarks of LLMs since I also recognise a lot of them but not so much here

Post reply on HN