feat: Enhance local command handling and introduce local intents
- Refactor `run_local_command` to manage subprocesses more effectively, ensuring child processes are terminated on timeout. - Introduce `_terminate` function to handle process group termination and capture output. - Implement `_command_output` to format command results with a character limit. - Add local intent recognition in `intents.py` to handle commands like "stop", "go to sleep", and "come here" without server interaction. - Normalize user input to match local intents while stripping filler words. - Update tests to cover new local intent functionality and ensure proper command handling. - Enhance speech processing to handle abbreviations and improve spoken output clarity.
This commit is contained in:
@@ -37,6 +37,42 @@ def rms(frame: np.ndarray) -> float:
|
||||
return float(np.sqrt(np.mean(frame.astype(np.float64) ** 2)))
|
||||
|
||||
|
||||
def flush(stream, max_seconds: float = 10.0, sample_rate: int = config.SAMPLE_RATE) -> int:
|
||||
"""Throw away whatever is already sitting in the mic's buffer. Returns the
|
||||
number of frames dropped.
|
||||
|
||||
PortAudio keeps capturing into a ring buffer while nothing is reading it, so
|
||||
audio recorded during a long blocking stretch is still queued when the next
|
||||
read happens. That matters exactly once: at the end of a reply the pet is
|
||||
about to listen for an answer, and the last fraction of a second of its own
|
||||
TTS is in that buffer. It's above the VAD threshold, so `record_utterance`
|
||||
treats it as the start of your answer, Deepgram transcribes it, and Bolt is
|
||||
handed his own sentence as if you had said it. With barge-in on, the
|
||||
detector was draining the stream during playback and the window is small;
|
||||
with `BARGE_IN=false` nothing drains it at all.
|
||||
|
||||
Only safe where the buffer is known to hold *nothing you said* — never
|
||||
before a wake-triggered recording, where the rest of "thunderbolt, what
|
||||
time is it" is legitimately queued and dropping it clips the request.
|
||||
|
||||
*max_seconds* bounds a single call so this can't chase a stream that's
|
||||
filling as fast as it's read. Best-effort: a fake stream in tests has no
|
||||
`read_available` and this is a no-op, which is the correct behaviour for
|
||||
one."""
|
||||
try:
|
||||
available = int(getattr(stream, "read_available", 0) or 0)
|
||||
except (TypeError, ValueError):
|
||||
return 0
|
||||
if available <= 0:
|
||||
return 0
|
||||
frames = min(available, int(max_seconds * sample_rate))
|
||||
try:
|
||||
stream.read(frames)
|
||||
except Exception:
|
||||
return 0 # a mid-flush device error is the reader's problem, not ours
|
||||
return frames
|
||||
|
||||
|
||||
def record_utterance(
|
||||
stream: AudioStream,
|
||||
should_continue=lambda: True,
|
||||
|
||||
Reference in New Issue
Block a user