"speech-to-speech" is a project by Huggingface - an easily deployable local VAD-STT-LLM-TTS pipeline.
I now have Majel Barrett (the voice of the computer in ST:TNG) answering my questions about the upcoming family vacation in Croatia. This is the stuff of dreams from the 90s having come real.
The VAD/TTS/STT components run on one computer and the LLM on another - everything GPU accelerated. The response time is even shorter than the one in TNG.
This household have no listening microphones or indoor cameras for privacy reasons, but a completely local setup on a separate VLAN ... hmm. I need to think about this.
My Star Trek-ified fork of the repo can be found here. It's quite close to behaving and sounding exactly like TNG, but you'll have to supply the right voice clone and chimes audio yourself ;)
This turned out to be so useful that I'm now looking at wiring up the house with microphones (wakeword protects regular conversations and everything's local anyway) and dedicate a GPU in an always-on computer for this.
I've never understood how a trekkie can be against LLMs. We're literally living Star Trek tech.
@troed Or why not HAL 9000? 🍾