Multi-monitor jumps, screen OCR, and generated sprite art

petctl gains screen verbs: `jump` (1-based number, name, next/prev/
primary/other, or a direction resolved from real geometry), `monitors`,
and `read` for OCR of a monitor's contents.

- monitors.py: pure layout model + jump-target resolution. The monitor
  list is published by PetWindow from QGuiApplication.screens() over a
  queued signal, so the controller and window agree on what "monitor 2"
  means; xrandr and Qt order screens differently on the same machine.
- screen_text.py: pull-only OCR (mss capture + Tesseract/RapidOCR).
  Nothing captures unless the server asks, and the text rides back up
  the tool-result relay so Bolt can read a screen mid-turn. Both deps
  optional, soft-failing with a reason. SCREEN_TEXT=false removes it.
- Query verbs are answered in controller._handle_command rather than
  pet_actions.describe(), because their output is the point.
- scripts/generate_bolt_sprites.py draws every frame; walk/ is a
  side-view cycle stepped by distance travelled, not by the animation
  timer, so the planted paw tracks the window exactly. sprite.py loads
  it via EXTRA_ANIMATIONS keyed by name, with has() so callers can
  decline a placeholder blob.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-07-28 16:17:11 -06:00
parent ccb3aeb7ae
commit b121bbba17
50 changed files with 2460 additions and 49 deletions
+24
View File
@@ -113,6 +113,30 @@ ELEVENLABS_VOICE_ID=
# error?" has a referent. Text only — no screenshots leave the machine.
#SCREEN_CONTEXT=true
# ── Monitors (optional) ─────────────────────────────────────────────────────
# Tacks a one-line summary of your screen layout onto each utterance (how
# many, their sizes, which one the pet is standing on) so Bolt can decide to
# `petctl jump 2` without asking what you've got plugged in. Costs nothing —
# the list comes from the UI, nothing is probed per turn.
#MONITOR_CONTEXT=true
# ── Screen text / OCR (optional) ────────────────────────────────────────────
# Lets Bolt read what's actually on a monitor with `petctl read [n|here|all]`
# and use it in his reply. Pull-only — nothing is captured unless he asks,
# and every read is logged.
#
# Needs the extras from requirements.txt plus an OCR engine:
# pip install mss pytesseract && sudo apt install tesseract-ocr
# or, without sudo:
# pip install mss rapidocr-onnxruntime
# mss captures on X11/Windows/macOS but NOT Wayland.
#
# This sends the text of a whole screen to the server when used. That's not a
# new capability — the shell relay could already screenshot and OCR — but it's
# a far easier one to reach for. SCREEN_TEXT=false removes it entirely.
#SCREEN_TEXT=true
#SCREEN_TEXT_MAX_CHARS=4000
# ── Quiet hours / do-not-disturb (optional) ────────────────────────────────
# Comma-separated HH:MM-HH:MM ranges; wrapping past midnight is fine. While
# napping the pet dims, stops wandering and makes no proactive noise — the