Files
Bolt-Pet/README.md
T

211 lines
11 KiB
Markdown

# Bolt Desktop Pet
A little animated pet that lives on your desktop and is just a face on top
of your Bolt server — same brain, memory, tools, and persona as Discord
chat and the Linux desk client. It talks to `ai/desk_api.py` on the server
exactly the way `desk_client/bolt_desk.py` does; this project only adds the
on-screen pet and swaps Deepgram/ElevenLabs playback to be cross-platform
(no `mpv`/`ffplay`/`espeak-ng` subprocess calls — pure `sounddevice`).
```
mic → wake-phrase spotter ("thunderbolt") / hotkey / click → record utterance
→ Deepgram STT (+ the focused window's title, for "what's this error?")
→ POST /desk/converse on your Bolt server → [server may relay a shell
command back to run on THIS machine, or a `petctl` command that moves
or emotes the pet] → reply → ElevenLabs streaming TTS → speakers
→ shown in a speech bubble + the pet's sprite state (idle/listening/
thinking/talking) updates the whole time
```
Nothing is sent to the server until the wake phrase fires, you press the
push-to-talk hotkey, or you click the pet — plus, if you turn them on, the
heartbeat and the desktop-notification bridge.
## Why a separate project instead of living in the tmn-api repo
This runs on your desktop machine, not the server — same relationship as
`desk_client/` (Linux) or the Android app, both of which are just clients of
the desk API over HTTP. It has no import dependency on the server repo at
all, so it can be copied anywhere and configured with its own `.env`.
## Setup
1. Copy this whole `bolt-pet/` folder to the machine you want the pet to run
on (if that isn't already this machine).
2. `cp .env.example .env` and fill in:
- `BOLT_SERVER_URL` + `DESK_API_KEY` — same as `desk_client/.env` on the
server side. Use the server's master `DESK_API_KEY`, or mint yourself a
personal one via the desk-only `api_key_generate` marker (see the main
repo's `CLAUDE.md` → "Per-user API keys").
- `DEEPGRAM_API_KEY` for STT.
- `ELEVENLABS_API_KEY` + `ELEVENLABS_VOICE_ID` for TTS (optional — falls
back to offline TTS via `pyttsx3` if omitted or if a request fails).
3. Run it:
- macOS/Linux: `./run.sh`
- Windows: `run.bat`
Both scripts create a local `.venv` and install `requirements.txt` on
first run. On Linux you'll also need system packages for audio:
`sudo apt install libportaudio2 espeak-ng`.
The pet appears near the bottom-right of your screen. It wanders off on its
own now and then; drag it anywhere and it tucks itself flush against a nearby
screen edge. That position isn't saved across restarts (see Known
limitations).
## Talking to it
- Say **"thunderbolt"** — detected fully on-device by a custom-trained
openWakeWord model (`thunderbolt.onnx`, ships in the project root), the
same way the server repo's `desk_client/bolt_desk.py` detects "hey bolt"
with `bolt.onnx`. Matches the `DEFAULT_WAKE_WORD` already used for Bolt's
Discord voice channels, so it's the same word everywhere.
- Or press **Ctrl+Alt+Space** (`PUSH_TO_TALK_HOTKEY`) from anywhere — useful
in a noisy room where the wake word misfires. Needs `pynput` and a session
that allows global key hooks; most Wayland sessions don't, in which case it
logs why at startup and everything else still works.
- Or just **click the pet** once (a drag doesn't count as a click).
- **Talk over it** to cut a long answer short — the mic stays live while it
speaks, and barging in starts your next turn immediately (`BARGE_IN`).
- Right-click the tray icon for **Talk now**, **Mute mic**, **Nap**,
**Wander around**, **Click through the pet**, **History…**, **Wake word
tuning…**, **Use default voice** and **Quit** — the pet window itself has
no title bar or taskbar entry.
- **Click the speech bubble** to copy what it just said; the tray's
**History…** window keeps the last `HISTORY_LIMIT` turns.
## What it does on its own
- **Wanders** the desktop while idle (`PET_WANDER`), stands still while
listening/thinking/talking or while a bubble is up.
- **Moves and emotes on command.** Bolt can relay `petctl move top-left`,
`petctl emote wave|hop|spin|nod|shake`, `petctl say ...`, `petctl wander
on|off`, `petctl nap on|off`, `petctl voice reset`, and `dialoguectl` for a
multi-voice scene. These are intercepted here and never reach a shell.
- **Naps** during `QUIET_HOURS` (e.g. `23:00-08:00`) or while a fullscreen
app is focused (`DND_ON_FULLSCREEN`) — it dims, stops wandering, and makes
no proactive noise. It still answers when you speak to it.
- **Reacts to desktop notifications** if you turn on `NOTIFICATION_BRIDGE`
(Linux/D-Bus) and set a `NOTIFICATION_FILTER` regex — matching
notifications get forwarded to the server, so it can tell you the deploy
went green. Off by default: each one costs a round trip.
## Speaking in another voice
Ask for a different voice — "use a clearer voice", "talk like a pirate", "say
that in Japanese" — and Bolt searches the ElevenLabs voice library on the
server, picks one, and tags his reply with it (`speak_as`); the pet is what
actually speaks in it. A Voice Library pick is added to your ElevenLabs
account automatically the first time it's used, and non-English replies (or
any picked voice) go through `ELEVENLABS_MULTILINGUAL_MODEL_ID` rather than
the English-only `eleven_flash_v2` default.
The new voice **stays on** for the rest of the conversation, because the
server tags a single reply and doesn't remember which voice it chose — so
"keep talking like that" would otherwise send it hunting for a voice again.
To get his own voice back: ask him ("use your normal voice" — he relays
`petctl voice reset`), use **Use default voice** in the tray menu (greyed
out unless a picked voice is active), or restart the pet. Set
`VOICE_STICKY=false` in `.env` if you'd rather each pick lasted exactly one
reply.
## Multi-voice dialogue
Ask for a scene — "do the argument between the two of them", "read that back
as a radio play" — and Bolt can relay a `dialoguectl` command that the pet
renders through ElevenLabs' Text to Dialogue endpoint: several voices in one
take, with delivery tags the v3 model acts on (`[cheerfully]`, `[whispering]`,
`[stuttering]`). One request per scene, so the voices actually react to each
other instead of sounding like clips glued together.
Name the cast in `.env` (`DIALOGUE_VOICES=narrator:9BWts…,villain:IKne3…`);
the name `self` always means whatever voice the pet is currently using, so
Bolt sounds like himself in his own scenes — including after a `speak_as`
switch. Scenes show up in the speech bubble with the tags stripped, count as
normal speech for the transcript, and can be talked over like any other reply.
`DIALOGUE=false` turns the whole thing off on this device.
## Wake-word detection
`bolt_pet/audio/wake_word.py` feeds every mic frame into `thunderbolt.onnx`
via the openWakeWord runtime (ONNX inference) and treats any class score at
or above `WAKE_WORD_THRESHOLD` (default `0.5`) as a detection — the exact
same per-frame `predict()`/`reset()` pattern as `desk_client/bolt_desk.py`'s
main loop. Point `WAKE_MODEL_FILE` in `.env` at a different `.onnx` model to
change the wake phrase later without touching any other code.
If it keeps ignoring you (or firing at the TV), open **Wake word tuning…**
from the tray: it shows the peak score while you talk and a rolling list of
near misses — frames that scored just under the threshold — and the slider
takes effect immediately, mid-listen. Set the threshold just below the peak
you can hit reliably, then write it into `.env` as `WAKE_WORD_THRESHOLD`.
## Project layout
```
bolt_pet/
config.py .env loading (same pattern as desk_client/bolt_desk.py)
state.py PetState enum + a small transition-checked state machine
server_client.py /desk/converse, /desk/tool_result, /desk/report_status
controller.py the pipeline: wake word -> STT -> server -> TTS, on a QThread
speech_text.py strips markdown/emoji/URLs so the voice never says "asterisk"
pet_actions.py petctl move/emote/say/wander/nap parsing
screen_context.py active-window title + fullscreen detection
quiet.py quiet-hours schedule
notifications.py desktop notification bridge (Linux/D-Bus)
history.py rolling conversation transcript
hotkey.py global push-to-talk (pynput, optional)
audio/
mic.py input stream + energy-based VAD utterance capture
wake_word.py openWakeWord thunderbolt.onnx detection (see above)
stt.py Deepgram
tts.py ElevenLabs streaming PCM, offline pyttsx3 fallback
barge_in.py "you started talking" detector, to cut playback short
ui/
app.py wires QApplication + window + tray + controller thread together
pet_window.py frameless/translucent/always-on-top sprite window + speech bubble
sprite.py frame animation loader (see assets/sprites/README.md)
tray.py system tray menu
history_window.py conversation scrollback (copyable)
wake_tuner.py live wake-word threshold + near-miss log
assets/sprites/ Kenney robot-pack art (CC0) — see assets/sprites/README.md
scripts/
slice_spritesheet.py cuts a grid sprite sheet into the per-frame convention
tests/ pure-logic unit tests (state machine, wake-phrase
matching, HTTP client against mocks) — nothing here
needs real audio hardware or a display
```
## Security notes
Same as `desk_client/bolt_desk.py`: the server can relay a shell command
back to this machine ("full desktop control" — "open firefox", "how full is
my disk", etc.), which this client executes as your desktop user with a
30-second timeout (`COMMAND_TIMEOUT_SECONDS`). That's the same trust model
as the Linux desk client and the Android app — commands only ever originate
from your own voice/click requests in your own session. Keep `DESK_API_KEY`
private; don't expose the desk API port to the open internet.
`petctl` commands (move/emote/say/wander/nap) are handled inside the pet and
never reach a shell, so that channel can't run anything.
Two features widen what leaves this machine, both off-switchable in `.env`:
`SCREEN_CONTEXT=true` (default) appends the focused window's *title* to what
you say — titles often contain file paths, document names or email subjects —
and `NOTIFICATION_BRIDGE=false` (default) can forward matching desktop
notifications. No screenshots or images are ever sent.
## Known limitations / not-yet-done
- Pet screen position isn't persisted across restarts.
- Wandering is a straight walk to a random point — no Shimeji-style physics,
wall-climbing or falling.
- Push-to-talk and the notification bridge are platform-limited: the hotkey
needs a session that allows global key hooks (most Wayland setups don't),
and the notification bridge is Linux/D-Bus only.
- Barge-in listens through the same mic that hears the pet's own voice. It
wants headphones or a decent gap between speaker and mic; if playback
interrupts itself, raise `BARGE_IN_RMS_THRESHOLD` or set `BARGE_IN=false`.
- Screen context is the window *title* only — the desk API takes text, so
there's no screenshot understanding.