Live data from Hacker News

My Journey to a reliable and enjoyable locally hosted voice assistant (2025)

community.home-assistant.io

151–153 of 153 posts

Re: My Journey to a reliable and enjoyable locally hosted voice assistant (2025)

#151
post #75

Earlier quoted context omitted.

You are reporting on a deliberately curated effort vs. what I understand is effectively voluntary data donation without incentives. It's not surprising to me that the later dataset ends up biased due to the differences in sourcing.

Your understanding of the datasets I helped create seems at odds with my experience actually creating the datasets. Do you have some insider experience or knowledge with dataset curation and creation for voice assistants that contradicts my own. The guideline is that the newer your model, the more likely it is to have diverse voice recognition datasets since it solves the earlier problems caused by non representative…

You're missing the point. No one cares about the datasets you've created in a commercial context.

The effort being discussed is a volunteer effort among a community of tech enthusiasts, who are disproportionately privacy-oriented vs the average person. This will undoubtable skew towards middle-aged male audiences, and will be extra-selective against children. It's a best-effort collection, they're probably not turning anyone away, and it's only what they can get, they're (AFAIK) not paying anyone to collect underrepresented demographics.

Re: My Journey to a reliable and enjoyable locally hosted voice assistant (2025)

#152

Earlier quoted context omitted.

How the hell are you managing that. On a simple micro level I can figure it out, but as a whole over the entire house? o.0

Home Assistant with adaptive lighting. https://github.com/basnijholt/adaptive-lighting

Oh cool, thanks.

Re: My Journey to a reliable and enjoyable locally hosted voice assistant (2025)

#153

Earlier quoted context omitted.

Your understanding of the datasets I helped create seems at odds with my experience actually creating the datasets. Do you have some insider experience or knowledge with dataset curation and creation for voice assistants that contradicts my own. The guideline is that the newer your model, the more likely it is to have diverse voice recognition datasets since it solves the earlier problems caused by non representative…

You're missing the point. No one cares about the datasets you've created in a commercial context. The effort being discussed is a volunteer effort among a community of tech enthusiasts, who are disproportionately privacy-oriented vs the average person. This will undoubtable skew towards middle-aged male audiences, and will be extra-selective against children. It's a best-effort collection, they're probably not turnin…

Ah.

I thought you were talking about voice assistants in general. My mustake

Post reply on HN