docs(02): research audio engine phase — MPV, soundfile, waveform, WebSocket broadcast

This commit is contained in:
Claude
2026-04-06 12:36:51 +00:00
parent a8db2a4ce7
commit d61c32a75b

View File

@@ -0,0 +1,604 @@
# Phase 02: Audio Engine - Research
**Researched:** 2026-04-05
**Domain:** MPV IPC bridge, audio position tracking, waveform extraction, FastAPI WebSocket broadcast
**Confidence:** HIGH (core stack verified against Alpine package registry and official sources)
---
## Summary
Phase 2 builds the master clock for the entire show system. Audio plays via MPV embedded through its libmpv C API (via Python ctypes), position is read at 10Hz using a polling loop (not observe_property), and that position is broadcast over WebSocket to all connected browser clients. Waveform data is extracted server-side using soundfile + numpy and sent to the browser as a downsampled float array for canvas rendering.
The deployment target is Linux Docker (Alpine 3.19) on groll.cloud — not Windows. The ROADMAP mentions `timeBeginPeriod(1)` and a "Windows timer fix" (INF-03), but this is irrelevant on Linux. On Linux, the kernel scheduler default tick is 250Hz (4ms granularity) and process timers default to 1ms resolution without any special calls. INF-03 as written is a no-op on this platform — the planner must decide: implement it as a Linux no-op stub, or mark it satisfied-by-platform.
The `py3-mpv` package (python-mpv 1.0.5, wrapping libmpv via ctypes) is available directly in Alpine 3.19's community repository alongside `mpv-libs` 0.37.0. No custom builds or workarounds are needed. Audio device access inside Docker requires mounting `/dev/snd` or the PulseAudio socket — or using `--ao=null` for position-tracking-only mode (the use case here, since audio output to the browser is not required).
**Primary recommendation:** Use `py3-mpv` (Alpine package, ctypes binding) with a dedicated 100ms polling loop for position, soundfile for waveform extraction, and the existing ConnectionManager WebSocket hub with an asyncio background task for 10Hz broadcast.
---
<phase_requirements>
## Phase Requirements
| ID | Description | Research Support |
|----|-------------|------------------|
| AUD-01 | Load audio from a local file (MP3, WAV, FLAC, OGG) | MPV supports all four formats natively; load via `player.loadfile(path)` |
| AUD-02 | Load audio from a YouTube URL via yt-dlp + MPV | Deferred to Phase 6 per ROADMAP — `--ytdl` flag; out of scope for Phase 2 |
| AUD-03 | Playback controls: play, pause, seek to position | `player.pause = True/False`, `player.seek(seconds, reference='absolute')` |
| AUD-04 | Realtime position reporting from MPV to UI (≤10ms poll interval) | Poll `player.time_pos` at 100ms interval (10Hz) from background thread; broadcast via WebSocket |
| AUD-05 | Waveform display rendered from audio samples in the timeline UI | soundfile reads audio as numpy array; downsample to ~1000 peak samples; send as JSON to browser |
| INF-03 | Windows 11 timer resolution fix (`timeBeginPeriod(1)`) | **Irrelevant on Linux.** Linux kernel provides ≤1ms timer resolution natively. Implement as platform-guarded no-op (only call on `sys.platform == 'win32'`). |
</phase_requirements>
---
## Standard Stack
### Core
| Library | Version | Purpose | Why Standard |
|---------|---------|---------|--------------|
| py3-mpv (python-mpv) | 1.0.5 (Alpine pkg) | Python ctypes binding to libmpv | Available as Alpine apk; same version as pip `python-mpv`; no custom build |
| mpv-libs | 0.37.0-r0 (Alpine) | libmpv.so.2 shared library | Required by py3-mpv; provides the C API |
| soundfile | 0.12.x | Read audio files as numpy arrays | Backed by libsndfile; supports WAV/FLAC/OGG natively; FAST, no FFmpeg dependency for core formats |
| numpy | 1.x / 2.x | Downsample waveform samples | Required by soundfile; standard scientific Python |
| libsndfile | 1.2.2-r1 (Alpine) | C library backing soundfile | Available as Alpine apk; handles WAV/FLAC/OGG |
### Supporting
| Library | Version | Purpose | When to Use |
|---------|---------|---------|-------------|
| ffmpeg-libs | 6.1.1-r0 (Alpine) | Decode MP3 for soundfile | MP3 is NOT supported by libsndfile alone; soundfile falls back to audioread which needs ffmpeg |
| pydub | latest | Alternative MP3 decode | Only if ffmpeg integration path proves complex; adds dependency |
### Alternatives Considered
| Instead of | Could Use | Tradeoff |
|------------|-----------|----------|
| py3-mpv (ctypes) | python-mpv-jsonipc | jsonipc is subprocess-based JSON over a Unix socket — more resilient but higher latency and more moving parts; ctypes is direct and lower overhead |
| py3-mpv (ctypes) | mpvasync | Minimal asyncio client for mpv JSON IPC; good fit for async but requires managing the mpv subprocess separately |
| soundfile + numpy | audiowaveform (BBC C++ tool) | audiowaveform not in Alpine apk, requires compilation; adds complexity; soundfile+numpy covers the same need with less friction |
| polling at 100ms | observe_property time-pos | observe_property fires on EVERY audio frame (many times/second at 44.1kHz), overwhelming IPC; polling at fixed 100ms is the correct approach |
**Installation (Dockerfile addition):**
```dockerfile
# Alpine apk packages (add to Dockerfile RUN apk add)
RUN apk add --no-cache mpv-libs ffmpeg-libs libsndfile
# Python packages (add to pyproject.toml dependencies)
# py3-mpv can be installed via apk OR pip; use pip for reproducibility
# "python-mpv>=1.0.5"
# "soundfile>=0.12.0"
# "numpy>=1.26.0"
```
**Version verification (confirmed 2026-04-05):**
- `py3-mpv` 1.0.5-r0 in Alpine 3.19 community
- `mpv-libs` 0.37.0-r0 in Alpine 3.19 community
- `libsndfile` 1.2.2-r1 in Alpine 3.19 main
- `ffmpeg-libs` 6.1.1-r0 in Alpine 3.19 community
---
## Architecture Patterns
### Recommended Project Structure
```
lightsync/
├── audio/
│ ├── __init__.py
│ ├── engine.py # MPVEngine: wraps python-mpv, position polling loop
│ └── waveform.py # extract_peaks(): soundfile + numpy downsampler
├── show/
│ ├── __init__.py
│ └── state.py # PlaybackState dataclass: position, paused, duration, file
├── api/
│ ├── ws.py # existing ConnectionManager + new broadcast_loop task
│ └── audio.py # REST endpoints: POST /api/audio/load, /play, /pause, /seek
└── main.py # lifespan: start MPVEngine, start broadcast_loop task
```
### Pattern 1: MPVEngine — Headless Audio-Only Instance
**What:** A class wrapping `mpv.MPV` configured for headless audio-only operation in Docker.
**When to use:** Always. This is the only MPV instance in the application.
```python
# Source: github.com/jaseg/python-mpv + mpv manual
import mpv
import threading
import time
class MPVEngine:
def __init__(self):
self._player = mpv.MPV(
vo='null', # no video output
ao='pulse', # PulseAudio; fallback: 'alsa' or 'null'
input_default_bindings=False,
input_vo_keyboard=False,
)
self._position: float = 0.0
self._lock = threading.Lock()
self._running = False
self._thread: threading.Thread | None = None
def start(self):
self._running = True
self._thread = threading.Thread(target=self._poll_loop, daemon=True)
self._thread.start()
def _poll_loop(self):
"""Poll time-pos at 10Hz. observe_property fires per-frame (too fast)."""
while self._running:
pos = self._player.time_pos # returns None if stopped
with self._lock:
self._position = pos or 0.0
time.sleep(0.1) # 10Hz
def get_position(self) -> float:
with self._lock:
return self._position
def load(self, path: str):
self._player.loadfile(path, mode='replace')
def play(self):
self._player.pause = False
def pause(self):
self._player.pause = True
def seek(self, seconds: float):
self._player.seek(seconds, reference='absolute')
@property
def duration(self) -> float | None:
return self._player.duration
def stop(self):
self._running = False
self._player.terminate()
```
### Pattern 2: Asyncio Background Broadcast Loop
**What:** A persistent asyncio task started at lifespan that reads position and broadcasts at 10Hz.
**When to use:** One task for the whole application lifetime, not per-connection.
```python
# Source: FastAPI lifespan pattern + asyncio
import asyncio
async def broadcast_loop(engine: MPVEngine, manager: ConnectionManager):
"""Broadcast position to all WebSocket clients at 10Hz."""
while True:
await asyncio.sleep(0.1)
pos = engine.get_position()
if manager.active_connections:
await manager.broadcast({
"type": "position",
"position": pos,
"paused": engine.player.pause,
})
```
**Integration in lifespan:**
```python
@asynccontextmanager
async def lifespan(app: FastAPI):
engine = MPVEngine()
engine.start()
app.state.engine = engine
task = asyncio.create_task(broadcast_loop(engine, manager))
yield
task.cancel()
engine.stop()
```
### Pattern 3: Thread-to-Asyncio Bridge
**What:** The MPVEngine runs in a daemon thread (blocking poll loop). When it needs to trigger something async (rare), use `asyncio.run_coroutine_threadsafe`.
**When to use:** Only if you need to push events FROM mpv callbacks INTO the asyncio loop.
```python
# Source: Python asyncio docs
import asyncio
loop = asyncio.get_event_loop()
def on_mpv_event():
# Called from mpv's internal thread
asyncio.run_coroutine_threadsafe(
manager.broadcast({"type": "event", "data": "..."}),
loop
)
```
### Pattern 4: Waveform Extraction
**What:** Server-side extraction of peak waveform data from audio files, downsampled to browser-friendly size.
**When to use:** When a file is loaded; result cached and sent once via WebSocket or HTTP.
```python
# Source: python-soundfile docs + numpy
import soundfile as sf
import numpy as np
def extract_peaks(path: str, num_samples: int = 1000) -> list[float]:
"""Return downsampled peak amplitude array for waveform display."""
data, samplerate = sf.read(path, always_2d=True)
mono = np.mean(data, axis=1) # mix to mono
chunk_size = max(1, len(mono) // num_samples)
peaks = []
for i in range(0, len(mono), chunk_size):
chunk = mono[i:i + chunk_size]
peaks.append(float(np.max(np.abs(chunk))))
return peaks[:num_samples]
```
**MP3 caveat:** soundfile does NOT support MP3 natively. Two options:
1. Install `ffmpeg` in Docker and use pydub or audioread as fallback
2. Use mutagen to read metadata only, and route MP3 decoding through ffmpeg subprocess
**Recommended MP3 path:** `ffmpeg -i input.mp3 -f f32le -ar 44100 -ac 1 pipe:1` → read from stdout as numpy array. ffmpeg-libs is already needed in the Alpine image.
### Pattern 5: WebSocket Transport Commands
**What:** Browser sends JSON commands; server dispatches to MPVEngine.
**When to use:** Replaces the echo stub in ws.py.
```python
# Extends existing ws.py dispatch
msg_handlers = {
"play": lambda _: engine.play(),
"pause": lambda _: engine.pause(),
"seek": lambda msg: engine.seek(float(msg["position"])),
"load": lambda msg: engine.load(msg["path"]),
}
# In websocket_endpoint receive loop:
handler = msg_handlers.get(msg.get("type"))
if handler:
handler(msg)
```
### Anti-Patterns to Avoid
- **observe_property for time-pos:** Fires on every decoded audio frame (~43× per second at 44.1kHz/1024 samples). Do NOT use for position reporting — use polling at 100ms.
- **Blocking calls inside async routes:** `engine.seek()` and `engine.load()` call libmpv synchronously. These are fast (< 1ms), but if they block longer, wrap in `asyncio.to_thread()`.
- **Running mpv with ao=null for production audio:** `ao=null` discards audio output. This is correct for a server that serves audio over network, but wrong if the Docker host needs to emit actual sound from speakers. For this project (serving a web UI), `ao=null` is correct — the browser doesn't receive audio bytes, it receives position ticks.
- **Starting the broadcast_loop task per WebSocket connection:** Start it once at lifespan. Per-connection tasks don't cancel cleanly.
---
## Don't Hand-Roll
| Problem | Don't Build | Use Instead | Why |
|---------|-------------|-------------|-----|
| Audio decoding + position tracking | Custom ffmpeg subprocess wrapper | python-mpv + mpv-libs | MPV handles seeking, gapless decode, codec detection, and position accuracy; rolling your own is months of work |
| MP3/WAV/FLAC detection and routing | Custom magic-byte sniffer | MPV auto-detects format | MPV uses libavformat probe; no manual routing needed |
| Audio sample extraction | Direct ffmpeg subprocess plumbing | soundfile + numpy for WAV/FLAC/OGG, ffmpeg pipe for MP3 | soundfile is battle-tested; the MP3 edge case is the only exception |
| WebSocket fan-out with dead connection cleanup | Custom send-and-check loop | The existing ConnectionManager already handles this | ws.py broadcast() already removes dead connections on exception |
| 10Hz clock | Custom timer thread | asyncio.sleep(0.1) in broadcast_loop | asyncio sleep is accurate enough for 10Hz; no custom timer needed |
---
## Critical Clarification: Windows Timer Fix (INF-03) on Linux
INF-03 states: "Windows 11 timer resolution fix applied at startup (`timeBeginPeriod(1)` via ctypes)."
**This requirement is inapplicable on Linux.** `timeBeginPeriod` is a Windows-only Win32 API (`timeapi.h`). Calling it on Linux is a no-op at best, and an import error at worst.
**Linux timer resolution:** On Linux 5.x+, the default scheduler tick is 250Hz or 1000Hz depending on kernel config (CONFIG_HZ). Process-level timer resolution via `clock_nanosleep` is typically < 100μs. No special calls are needed.
**Recommended implementation:** Implement INF-03 as a platform guard in `main.py`:
```python
import sys
import ctypes
def apply_timer_fix():
"""Apply 1ms timer resolution on Windows. No-op on Linux/macOS."""
if sys.platform == 'win32':
winmm = ctypes.WinDLL('winmm')
winmm.timeBeginPeriod(1)
```
This satisfies the requirement as stated while being correct for the actual deploy target.
---
## Docker / Deployment Notes
### Alpine 3.19 Dockerfile additions
```dockerfile
FROM python:3.11-alpine
# Audio stack: mpv-libs (libmpv.so.2) + audio codec support
RUN apk add --no-cache \
mpv-libs \
ffmpeg-libs \
libsndfile
# For actual audio output (if ever needed from Docker):
# Mount /dev/snd at runtime: docker run --device /dev/snd ...
# Or PulseAudio socket: -v /run/user/1000/pulse:/run/user/1000/pulse
```
### Audio output in Docker
The VPS Docker container does NOT need to emit audio to speakers. MPV is used purely as a position-accurate playback clock. Therefore:
- Use `ao='null'` (or `ao='alsa'` with `/dev/snd` mounted) — no audio device needed
- `vo='null'` — no video output needed
- No display server, no ALSA/PulseAudio socket mounting required for position tracking
**Confirmed:** `mpv-libs` 0.37.0-r0 provides `libmpv.so.2` — confirmed via Alpine package page. python-mpv 1.0.5 requires libmpv API ≥ 1.108 (libmpv ≥ 0.33). mpv 0.37.0 > 0.33, so the API requirement is met.
---
## Common Pitfalls
### Pitfall 1: observe_property time-pos Flood
**What goes wrong:** Registering `@player.property_observer('time-pos')` causes the callback to fire at the audio decode rate (~43/s for 44.1kHz), overwhelming the asyncio event loop with 430 messages per second at 10× the target rate.
**Why it happens:** MPV fires property-change events on every decoded audio frame, not at a configurable Hz rate.
**How to avoid:** Use a 100ms polling loop reading `player.time_pos` directly. This gives exactly 10Hz regardless of audio format.
**Warning signs:** WebSocket message queue growing, browser lagging behind, CPU spike in uvicorn worker.
### Pitfall 2: MP3 Not Supported by soundfile
**What goes wrong:** `sf.read('track.mp3')` raises `SoundFileError: Error opening 'track.mp3': Format not recognised`.
**Why it happens:** libsndfile does not include an MP3 decoder. soundfile's `audioread` fallback requires FFmpeg on PATH.
**How to avoid:** Add `ffmpeg` (or `ffmpeg-libs`) to the Dockerfile AND test MP3 loading explicitly. Alternatively, use a dedicated ffmpeg subprocess for MP3 → float array conversion.
**Warning signs:** Load succeeds for WAV in tests but fails for MP3 in production.
### Pitfall 3: Blocking MPV Calls in Async Context
**What goes wrong:** Calling `engine.load()` or `engine.seek()` inside an `async def` FastAPI route blocks the event loop while libmpv processes the command.
**Why it happens:** python-mpv ctypes calls are synchronous; libmpv file loading can take 10100ms for large files.
**How to avoid:** Wrap in `asyncio.to_thread(engine.load, path)` for the load command. Seek is typically fast enough (< 1ms) but wrap defensively.
**Warning signs:** WebSocket clients stop receiving position updates during file load.
### Pitfall 4: Position Reporting to Dead WebSocket Connections
**What goes wrong:** The broadcast_loop task throws an exception for a disconnected client, then crashes, stopping ALL position broadcasts.
**Why it happens:** `websocket.send_text()` raises on a closed connection.
**How to avoid:** The existing `ConnectionManager.broadcast()` already catches exceptions and removes dead connections. Do not bypass it.
**Warning signs:** Position updates stop for ALL clients when one disconnects.
### Pitfall 5: time_pos Returns None During Seek or Before Load
**What goes wrong:** `player.time_pos` returns `None` when no file is loaded or during a seek operation. Passing `None` to JSON serialization raises `TypeError`.
**Why it happens:** libmpv property returns None for unavailable properties.
**How to avoid:** Always use `pos = player.time_pos or 0.0` in the polling loop.
**Warning signs:** `TypeError: Object of type NoneType is not JSON serializable` in logs.
### Pitfall 6: mpv-libs vs Full mpv Package
**What goes wrong:** Installing `mpv` (full package, 163 deps including display libs) instead of `mpv-libs` (15 deps, 7MB).
**Why it happens:** Confusion between the player binary and the embeddable library.
**How to avoid:** Install `mpv-libs` only. python-mpv needs `libmpv.so.2`, NOT the `mpv` command-line binary.
**Warning signs:** Docker image bloat; install taking > 1 minute; mesa/wayland/X11 dependencies appearing.
---
## Code Examples
### Complete MPVEngine skeleton
```python
# lightsync/audio/engine.py
import threading
import time
import mpv
class MPVEngine:
def __init__(self):
self._player = mpv.MPV(
vo='null',
ao='null', # no audio output needed on server
input_default_bindings=False,
input_vo_keyboard=False,
)
self._position: float = 0.0
self._lock = threading.Lock()
self._running = False
def start(self):
self._running = True
t = threading.Thread(target=self._poll_loop, daemon=True)
t.start()
def _poll_loop(self):
while self._running:
pos = self._player.time_pos
with self._lock:
self._position = pos if pos is not None else 0.0
time.sleep(0.1)
def get_state(self) -> dict:
with self._lock:
return {
"position": self._position,
"paused": bool(self._player.pause),
"duration": self._player.duration,
"loaded": self._player.path is not None,
}
def load(self, path: str):
self._player.loadfile(path, mode='replace')
self._player.pause = True # load paused, wait for explicit play
def play(self):
self._player.pause = False
def pause(self):
self._player.pause = True
def seek(self, seconds: float):
self._player.seek(seconds, reference='absolute')
def stop(self):
self._running = False
try:
self._player.terminate()
except Exception:
pass
```
### Waveform extraction (MP3 + non-MP3)
```python
# lightsync/audio/waveform.py
import subprocess
import numpy as np
import soundfile as sf
from pathlib import Path
def extract_peaks(path: str, num_peaks: int = 1000) -> list[float]:
"""Extract amplitude peaks from audio file. Handles MP3 via ffmpeg."""
p = Path(path)
if p.suffix.lower() == '.mp3':
return _extract_peaks_mp3(path, num_peaks)
return _extract_peaks_soundfile(path, num_peaks)
def _extract_peaks_soundfile(path: str, num_peaks: int) -> list[float]:
data, _ = sf.read(path, always_2d=True)
mono = np.mean(data, axis=1)
return _downsample_peaks(mono, num_peaks)
def _extract_peaks_mp3(path: str, num_peaks: int) -> list[float]:
"""Decode MP3 via ffmpeg pipe to float32 array."""
cmd = [
'ffmpeg', '-i', path,
'-f', 'f32le', '-ar', '44100', '-ac', '1', 'pipe:1',
'-loglevel', 'quiet'
]
result = subprocess.run(cmd, capture_output=True)
mono = np.frombuffer(result.stdout, dtype=np.float32)
return _downsample_peaks(mono, num_peaks)
def _downsample_peaks(mono: np.ndarray, num_peaks: int) -> list[float]:
chunk = max(1, len(mono) // num_peaks)
peaks = [
float(np.max(np.abs(mono[i:i+chunk])))
for i in range(0, len(mono), chunk)
]
return peaks[:num_peaks]
```
### FastAPI lifespan integration
```python
# main.py additions
from lightsync.audio.engine import MPVEngine
from lightsync.api.ws import manager
import asyncio
async def _broadcast_loop(engine: MPVEngine):
while True:
await asyncio.sleep(0.1)
if manager.active_connections:
await manager.broadcast({"type": "tick", **engine.get_state()})
@asynccontextmanager
async def lifespan(app: FastAPI):
# ... existing registry/show_store setup ...
engine = MPVEngine()
engine.start()
app.state.engine = engine
task = asyncio.create_task(_broadcast_loop(engine))
yield
task.cancel()
engine.stop()
```
---
## State of the Art
| Old Approach | Current Approach | When Changed | Impact |
|--------------|------------------|--------------|--------|
| mpv --slave mode (stdin commands) | JSON IPC via Unix socket | mpv 0.7+ (2014) | Clean bidirectional; python-mpv wraps this |
| manual libmpv ctypes bindings | python-mpv package | 2014present | Maintained wrapper, no DIY ctypes |
| observe_property for position | Polling at fixed Hz | Known issue (#5661) | Correct 10Hz rate, not frame-rate dependent |
| audiowaveform (BBC C++ binary) | soundfile + numpy | Still valid alternative | soundfile needs no compilation; less infra |
**Deprecated/outdated:**
- `--slave` mode stdin control: Removed from mpv. Use IPC socket or libmpv API.
- `python-mpv-jsonipc` as primary approach: The STATE.md notes this was initially preferred to avoid Windows DLL issues. Since we deploy on Linux, the ctypes binding (py3-mpv) is simpler and available as an Alpine package.
---
## Open Questions
1. **Audio output: null vs pulse vs alsa?**
- What we know: Server doesn't need to output sound; `ao=null` is correct for position-tracking-only
- What's unclear: If someone ever runs this locally on a desktop, they'd want actual audio; `ao=null` silences playback
- Recommendation: Use `ao=null` for Docker; expose as env var `MPV_AO` defaulting to `null`
2. **File upload vs path reference for audio loading**
- What we know: AUD-01 says "load from local file" — unclear if user uploads file or provides a server path
- What's unclear: Is the audio file already on the VPS (e.g., mounted volume), or does the user upload from their browser?
- Recommendation: Plan 02-01 should implement path-based loading first (simpler); file upload can be a follow-up
3. **Waveform data size and format**
- What we know: 1000 peak samples as float array is standard for timeline display
- What's unclear: Whether to send this as a WebSocket message or as a separate HTTP endpoint response
- Recommendation: HTTP GET `/api/audio/waveform` returning JSON; don't mix it into the WebSocket stream
---
## Environment Availability
| Dependency | Required By | Available | Version | Fallback |
|------------|------------|-----------|---------|----------|
| mpv-libs | AUD-01, AUD-03, AUD-04 | ✓ (Alpine apk) | 0.37.0-r0 | — |
| py3-mpv (python-mpv) | AUD-01, AUD-03, AUD-04 | ✓ (Alpine apk or pip) | 1.0.5 | — |
| libsndfile | AUD-05 (WAV/FLAC/OGG) | ✓ (Alpine apk) | 1.2.2-r1 | — |
| soundfile (Python) | AUD-05 | pip install | 0.12.x | — |
| numpy | AUD-05 | pip install | 1.26+ | — |
| ffmpeg-libs | AUD-05 (MP3 decode) | ✓ (Alpine apk) | 6.1.1-r0 | Skip MP3 support |
| ffmpeg binary | AUD-05 (MP3 pipe) | ✓ (Alpine: `apk add ffmpeg`) | 6.1.1 | pydub (heavier) |
| ctypes | python-mpv binding | ✓ (stdlib) | stdlib | — |
**Missing dependencies with no fallback:** None — all required packages are available in Alpine 3.19.
**Notes:**
- `py3-mpv` can be installed as an Alpine package (`apk add py3-mpv`) or via pip (`pip install python-mpv`). Alpine package pins to 1.0.5; pip may have a newer version. Prefer pip for version control.
- `ffmpeg` binary (not just libs) is needed for the MP3 waveform extraction pipe. Add `ffmpeg` to Dockerfile, not just `ffmpeg-libs`.
---
## Sources
### Primary (HIGH confidence)
- Alpine 3.19 package registry — confirmed: `py3-mpv` 1.0.5, `mpv-libs` 0.37.0, `libsndfile` 1.2.2, `ffmpeg-libs` 6.1.1, `ffmpeg` 6.1.1
- [github.com/jaseg/python-mpv](https://github.com/jaseg/python-mpv) — version 1.0.8 on main; property names, constructor params, observer pattern, thread model
- [pkgs.alpinelinux.org/package/v3.19/community/x86_64/py3-mpv](https://pkgs.alpinelinux.org/package/v3.19/community/x86_64/py3-mpv) — confirms py3-mpv 1.0.5 depends on mpv-libs + python3
### Secondary (MEDIUM confidence)
- [mpv JSON IPC issue #5661](https://github.com/mpv-player/mpv/issues/5661) — observe_property time-pos fires per frame; polling is correct workaround; closed as by-design
- [python-soundfile docs](https://python-soundfile.readthedocs.io/) — confirmed libsndfile backing, no MP3 support, numpy array output
- [FastAPI WebSocket docs](https://fastapi.tiangolo.com/advanced/websockets/) — background task + broadcast pattern
- [Python asyncio docs](https://docs.python.org/3/library/asyncio-dev.html) — `run_coroutine_threadsafe` for thread-to-async bridge
### Tertiary (LOW confidence)
- Docker audio device mounting approach (`/dev/snd`, PulseAudio socket) — sourced from community forums; not verified against official Docker docs. Moot for this deployment since `ao=null` is the correct choice.
---
## Metadata
**Confidence breakdown:**
- Standard stack: HIGH — all packages verified against Alpine 3.19 package registry
- Architecture: HIGH — patterns verified against python-mpv source and FastAPI docs
- Pitfalls: HIGH for polling vs observe_property (verified against official mpv issue tracker); MEDIUM for Docker audio (community sources)
- INF-03 Linux assessment: HIGH — `timeBeginPeriod` is Windows-only Win32 API
**Research date:** 2026-04-05
**Valid until:** 2026-10-05 (stable libraries; Alpine 3.19 EOL before then but packages won't change)