Multi-monitor jumps, screen OCR, and generated sprite art
petctl gains screen verbs: `jump` (1-based number, name, next/prev/ primary/other, or a direction resolved from real geometry), `monitors`, and `read` for OCR of a monitor's contents. - monitors.py: pure layout model + jump-target resolution. The monitor list is published by PetWindow from QGuiApplication.screens() over a queued signal, so the controller and window agree on what "monitor 2" means; xrandr and Qt order screens differently on the same machine. - screen_text.py: pull-only OCR (mss capture + Tesseract/RapidOCR). Nothing captures unless the server asks, and the text rides back up the tool-result relay so Bolt can read a screen mid-turn. Both deps optional, soft-failing with a reason. SCREEN_TEXT=false removes it. - Query verbs are answered in controller._handle_command rather than pet_actions.describe(), because their output is the point. - scripts/generate_bolt_sprites.py draws every frame; walk/ is a side-view cycle stepped by distance travelled, not by the animation timer, so the planted paw tracks the window exactly. sprite.py loads it via EXTRA_ANIMATIONS keyed by name, with has() so callers can decline a placeholder blob. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -11,13 +11,18 @@ dependency** on the server repo; it's a standalone HTTP client configured via
|
||||
its own `.env`.
|
||||
|
||||
Pipeline: `mic → openWakeWord ("thunderbolt", on-device) / push-to-talk /
|
||||
click → record utterance → Deepgram STT → + active-window context → POST
|
||||
/desk/converse → [server may relay a shell command to run on this machine, or
|
||||
a `petctl` pseudo-command that moves/emotes the pet instead] → reply →
|
||||
click → record utterance → Deepgram STT → + active-window + screen-layout
|
||||
context → POST /desk/converse → [server may relay a shell command to run on
|
||||
this machine, or a `petctl` pseudo-command that moves/emotes the pet, jumps it
|
||||
to another monitor, or reads a screen's text back instead] → reply →
|
||||
ElevenLabs streaming TTS (or offline pyttsx3 fallback) → speakers`, with the
|
||||
pet sprite/speech bubble reflecting state throughout, and playback
|
||||
interruptible by talking over it (barge-in).
|
||||
|
||||
Because a relayed command's output goes back up the tool-result relay before
|
||||
the final reply, a `petctl read` mid-turn means Bolt can look at a monitor and
|
||||
then talk about what's on it in the same answer.
|
||||
|
||||
Side channels that let the pet act between turns: the heartbeat (proactive
|
||||
announcements), the desktop notification bridge, and autonomous wandering —
|
||||
all suppressed while it's napping (quiet hours / fullscreen DND).
|
||||
@@ -34,6 +39,9 @@ run.bat # Windows
|
||||
QT_QPA_PLATFORM=offscreen .venv/bin/pytest tests/
|
||||
.venv/bin/pytest tests/test_state.py::test_happy_path_transitions # single test
|
||||
|
||||
# Redraw the pet's sprite frames (the committed PNGs are this script's output)
|
||||
python scripts/generate_bolt_sprites.py # --out /tmp/x to preview first
|
||||
|
||||
# Convert a grid sprite sheet into the per-frame-PNG convention sprite.py expects
|
||||
python scripts/slice_spritesheet.py path/to/sheet.png assets/sprites/idle --cols 6 --rows 1
|
||||
```
|
||||
@@ -90,17 +98,46 @@ logs a missing-config message and exits its thread instead of starting.
|
||||
accepts an injectable stream/model/protocol so tests don't need real audio
|
||||
hardware or a display.
|
||||
- **`pet_actions.py`** — `petctl` pseudo-commands (`petctl move top-left`,
|
||||
`petctl emote wave`, `say`/`wander`/`nap`). The desk API has no "move the
|
||||
pet" payload type and this repo can't change the server, so these ride the
|
||||
existing shell-command relay: `controller._handle_command` parses them and
|
||||
they never reach `subprocess`; anything else is a real shell command exactly
|
||||
as before. Pure parsing; the UI half is `PetWindow.apply_action`.
|
||||
`petctl emote wave`, `say`/`wander`/`nap`, plus the screen verbs
|
||||
`jump`/`monitors`/`read`). The desk API has no "move the pet" payload type
|
||||
and this repo can't change the server, so these ride the existing
|
||||
shell-command relay: `controller._handle_command` parses them and they never
|
||||
reach `subprocess`; anything else is a real shell command exactly as before.
|
||||
Pure parsing; the UI half is `PetWindow.apply_action`. Note `jump`'s target
|
||||
is *not* validated here — which monitors exist is a runtime fact this pure
|
||||
module doesn't have, so the spec passes through to `monitors.resolve()`.
|
||||
Query verbs (`monitors`, `read`) are answered in `_handle_command` rather
|
||||
than by `pet_actions.describe()`, because their output *is* the point: it
|
||||
goes back up the tool-result relay for Bolt to use in his reply.
|
||||
- **`screen_context.py`** — active-window title (xprop/xdotool, Win32,
|
||||
osascript) appended to each utterance via `context_for()`, plus
|
||||
`is_fullscreen_active()` for do-not-disturb. Text only — the desk API takes
|
||||
no images. Every probe is best-effort and returns None/False rather than
|
||||
raising; the parsing is split into pure functions that are tested without a
|
||||
display server.
|
||||
- **`monitors.py`** — the screen layout, and resolving `petctl jump` targets
|
||||
(a 1-based number, a name, `next`/`prev`/`primary`/`other`, or a direction
|
||||
like `left`/`up` worked out from the actual geometry). Pure — no Qt, no
|
||||
subprocess. The monitor list is *published by the UI*
|
||||
(`PetWindow.publish_monitors` builds it from `QGuiApplication.screens()` and
|
||||
emits it over a queued signal to `controller.set_monitors`), because the
|
||||
controller and the window must agree on what "monitor 2" means: enumerating
|
||||
with `xrandr` on one side and Qt's screen list on the other gives different
|
||||
orderings on the same machine, and Bolt would announce one screen and land
|
||||
on another. Qt is the single source of truth; `Monitor.index` is 0-based and
|
||||
`.number` is the 1-based value used in every string a human or the model
|
||||
sees. The controller resolves a jump to a concrete index *before* emitting
|
||||
it, so the window can't re-resolve against a different list.
|
||||
- **`screen_text.py`** — OCR, so Bolt can read what's on a monitor
|
||||
(`petctl read [n|here|all]`). **Pull, not push**: nothing captures on its own
|
||||
— the server has to ask, and the text goes back as that command's output.
|
||||
That's deliberate; OCR of a 4K screen costs a second or two that would
|
||||
otherwise be added to *every* utterance, and screen contents leaving the
|
||||
machine should be a visible decision rather than a constant. Capture needs
|
||||
`mss` (X11/Win32/macOS, **not** Wayland), recognition needs Tesseract or
|
||||
RapidOCR; both are optional and soft-fail with a reason the way `hotkey.py`
|
||||
does, and `read_monitor()` never raises because its return value is command
|
||||
output. Engine selection takes injected probes so it's testable wherever.
|
||||
- **`quiet.py`** — quiet-hours spec parsing (`23:00-08:00`, wraps midnight,
|
||||
comma-separated). Napping suppresses *proactive* noise and wandering only;
|
||||
wake word / click / push-to-talk still work.
|
||||
@@ -162,6 +199,15 @@ logs a missing-config message and exits its thread instead of starting.
|
||||
target every `PET_WANDER_INTERVAL_SECONDS` (randomized), suppressed
|
||||
whenever the pet is non-IDLE, napping, dragged, or has a bubble up. A
|
||||
commanded `petctl move` overrides all of that except the drag.
|
||||
- **the walk cycle** — while actually travelling, `_animation_key()` swaps
|
||||
the state animation for the side-view `walk/` frames (not a `PetState` —
|
||||
see the sprites README). It is stepped by *distance travelled*
|
||||
(`_WALK_PIXELS_PER_FRAME`), never by the animation timer, so the planted
|
||||
paw tracks backwards at exactly the speed the window moves forwards;
|
||||
`_advance_frame` deliberately no-ops while walking so the two can't
|
||||
double-step it. The art is drawn facing right and `_oriented()` mirrors it
|
||||
(cached per frame) when heading left. No `walk/` art → falls back to the
|
||||
old coded bob rather than a placeholder blob.
|
||||
- **emotes** — `emote_transform()` is pure maths (dx, dy, rotation, scale
|
||||
from a 0..1 progress) kept out of `paintEvent` so the curves are unit
|
||||
tested; every emote must return to the identity transform at progress 1.0
|
||||
@@ -174,9 +220,14 @@ logs a missing-config message and exits its thread instead of starting.
|
||||
- **edge snapping** (`PET_EDGE_SNAP`) after a drag or a stroll, and **nap
|
||||
dimming** (`set_napping`).
|
||||
`sprite.py` loads `assets/sprites/<state>/*.png` (filename-sorted, looping —
|
||||
currently Kenney's CC0 robot pack, see `assets/sprites/README.md`) and falls
|
||||
the art is *generated* by `scripts/generate_bolt_sprites.py`, a Pillow
|
||||
drawing of Bolt as a shepherd pup; edit the script and re-run it rather than
|
||||
the committed PNGs, see `assets/sprites/README.md`) and falls
|
||||
back to a procedurally-drawn placeholder blob per state if a folder has no
|
||||
frames; `tray.py` is the system tray menu (talk now / mute / nap / wander /
|
||||
frames. It also loads `EXTRA_ANIMATIONS` — currently just `walk/` — keyed by
|
||||
name rather than by `PetState`, with `has()` reporting whether a key is
|
||||
backed by real art so callers can decline a placeholder instead of trotting
|
||||
a blob across the desktop; `tray.py` is the system tray menu (talk now / mute / nap / wander /
|
||||
click-through / history / wake-word tuning / quit) — the pet window has no
|
||||
title bar or taskbar entry; `history_window.py` and `wake_tuner.py` are the
|
||||
two dialogs it opens.
|
||||
@@ -262,9 +313,23 @@ can't reach the user's PipeWire socket from a root session (raw ALSA devices
|
||||
reject the 16 kHz capture rate — `paInvalidSampleRate`), and every relayed
|
||||
command would run unconstrained.
|
||||
|
||||
Two newer features widen what leaves this machine, both switchable in `.env`:
|
||||
Several features widen what leaves this machine, all switchable in `.env`:
|
||||
`SCREEN_CONTEXT` appends the focused window's *title* to each utterance
|
||||
(titles often contain file paths, document names, or subject lines), and
|
||||
`NOTIFICATION_BRIDGE` (off by default) forwards matching desktop
|
||||
notifications to the server. Neither sends screenshots or notification
|
||||
contents you haven't matched with `NOTIFICATION_FILTER`.
|
||||
(titles often contain file paths, document names, or subject lines),
|
||||
`MONITOR_CONTEXT` appends the screen layout (sizes and names only — no
|
||||
contents), and `NOTIFICATION_BRIDGE` (off by default) forwards matching
|
||||
desktop notifications to the server. None of those send screenshots or
|
||||
notification contents you haven't matched with `NOTIFICATION_FILTER`.
|
||||
|
||||
`SCREEN_TEXT` is the biggest of them: `petctl read` OCRs a whole monitor and
|
||||
sends the recognised text to the server — everything visible, not just the
|
||||
focused window. Two things keep it honest. It's **pull-only**: no capture
|
||||
happens unless the server explicitly asks, so it can't leak in the background
|
||||
the way a per-turn annotation would, and each read is logged. And it is
|
||||
strictly *not* a new capability — the shell relay could already run a
|
||||
screenshot tool and pipe it through OCR — it just makes a thing the trust
|
||||
model already allowed reliable, bounded (`SCREEN_TEXT_MAX_CHARS`) and
|
||||
visible. It is nonetheless far easier to reach for than the shell route, so
|
||||
if that trade isn't one you want, `SCREEN_TEXT=false` removes it and
|
||||
`petctl read` starts reporting that it's disabled. Capture is `mss`-based and
|
||||
therefore silently unavailable on Wayland.
|
||||
|
||||
Reference in New Issue
Block a user