22 KiB
phase, plan, type, wave, depends_on, files_modified, autonomous, requirements, must_haves
| phase | plan | type | wave | depends_on | files_modified | autonomous | requirements | must_haves | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 02-audio-engine | 03 | execute | 3 |
|
|
false |
|
|
Purpose: The user needs visual feedback — waveform shows song structure, transport controls audio, position cursor tracks playback in real time. Output: Working transport bar with play/pause/seek, waveform canvas in timeline panel, real-time position updates from WebSocket ticks.
<execution_context> @$HOME/.claude/get-shit-done/workflows/execute-plan.md @$HOME/.claude/get-shit-done/templates/summary.md </execution_context>
@.planning/PROJECT.md @.planning/ROADMAP.md @.planning/STATE.md @.planning/phases/02-audio-engine/02-RESEARCH.md @.planning/phases/02-audio-engine/02-01-SUMMARY.md @.planning/phases/02-audio-engine/02-02-SUMMARY.md From lightsync/audio/engine.py: ```python class MPVEngine: def get_state(self) -> dict[str, Any]: # Returns: {"position": float, "paused": bool, "duration": float|None, "loaded": bool, "file": str|None} ```WebSocket tick message format:
{"type": "tick", "position": 42.1, "paused": false, "duration": 180.5, "loaded": true, "file": "/path/to/song.mp3"}
WebSocket commands accepted:
{"type": "play"}
{"type": "pause"}
{"type": "seek", "position": 42.0}
{"type": "load", "path": "/path/to/file.mp3"}
REST endpoints:
POST /api/audio/load — body: {"path": "..."}
GET /api/audio/state — returns engine state dict
From lightsync/api/audio.py:
router = APIRouter()
# Already has POST /load and GET /state
From lightsync/frontend/index.html:
<div id="timeline-panel">— currently shows placeholder text<footer class="transport-bar" id="transport-panel">— currently shows placeholder text<script src="/app.js" type="module">— ES modules
From lightsync/frontend/app.js:
class LightSyncClientwithhandleMessage(msg),send(msg),connect()const client = new LightSyncClient()— global instance- Uses
wss://protocol detection:${location.protocol === 'https:' ? 'wss' : 'ws'}://${location.host}/ws
"""Waveform peak extraction for timeline display (AUD-05)."""
import subprocess
from pathlib import Path
import numpy as np
import soundfile as sf
def extract_peaks(path: str, num_peaks: int = 1000) -> list[float]:
"""Extract downsampled peak amplitudes from audio file.
Returns a list of floats (0.0-1.0) representing peak amplitude per chunk.
Handles MP3 via ffmpeg subprocess (libsndfile doesn't support MP3).
Args:
path: Absolute path to audio file
num_peaks: Number of peak samples to return (default 1000)
Returns:
List of float peak values, length <= num_peaks
"""
p = Path(path)
if p.suffix.lower() == ".mp3":
return _extract_peaks_mp3(path, num_peaks)
return _extract_peaks_soundfile(path, num_peaks)
def _extract_peaks_soundfile(path: str, num_peaks: int) -> list[float]:
"""Extract peaks using soundfile (WAV, FLAC, OGG)."""
data, _ = sf.read(path, always_2d=True)
mono = np.mean(data, axis=1)
return _downsample_peaks(mono, num_peaks)
def _extract_peaks_mp3(path: str, num_peaks: int) -> list[float]:
"""Extract peaks from MP3 via ffmpeg pipe to float32 PCM."""
cmd = [
"ffmpeg", "-i", path,
"-f", "f32le", "-ar", "44100", "-ac", "1",
"pipe:1", "-loglevel", "quiet",
]
result = subprocess.run(cmd, capture_output=True, timeout=60)
if result.returncode != 0:
raise RuntimeError(f"ffmpeg failed for {path}: exit code {result.returncode}")
mono = np.frombuffer(result.stdout, dtype=np.float32)
return _downsample_peaks(mono, num_peaks)
def _downsample_peaks(mono: np.ndarray, num_peaks: int) -> list[float]:
"""Downsample mono audio to peak amplitude array."""
if len(mono) == 0:
return []
chunk_size = max(1, len(mono) // num_peaks)
peaks = []
for i in range(0, len(mono), chunk_size):
chunk = mono[i : i + chunk_size]
peaks.append(float(np.max(np.abs(chunk))))
return peaks[:num_peaks]
- Update
lightsync/api/audio.py— add waveform endpoint:
Add import at top: from lightsync.audio.waveform import extract_peaks
Add endpoint AFTER the existing ones:
@router.get("/waveform")
async def get_waveform(request: Request, peaks: int = 1000):
"""Get waveform peak data for the currently loaded audio file (AUD-05).
Query params:
peaks: Number of peak samples (default 1000, max 5000)
"""
engine = request.app.state.engine
if engine is None:
raise HTTPException(status_code=503, detail="Audio engine not ready")
state = engine.get_state()
if not state["loaded"] or not state["file"]:
raise HTTPException(status_code=400, detail="No audio file loaded")
num_peaks = min(max(peaks, 100), 5000)
try:
data = await asyncio.to_thread(extract_peaks, state["file"], num_peaks)
except Exception as e:
raise HTTPException(status_code=500, detail=f"Waveform extraction failed: {e}")
return {"peaks": data, "count": len(data), "file": state["file"]}
Uses asyncio.to_thread because waveform extraction can take seconds for large files.
cd /home/claude/led2 && test -f lightsync/audio/waveform.py && grep -q "def extract_peaks" lightsync/audio/waveform.py && grep -q "_extract_peaks_mp3" lightsync/audio/waveform.py && grep -q "_extract_peaks_soundfile" lightsync/audio/waveform.py && grep -q "waveform" lightsync/api/audio.py && grep -q "extract_peaks" lightsync/api/audio.py && echo "PASS"
<acceptance_criteria>
- File
lightsync/audio/waveform.pyexists withextract_peaks(path, num_peaks=1000) -> list[float] - MP3 files handled via ffmpeg subprocess pipe (
_extract_peaks_mp3) - WAV/FLAC/OGG handled via soundfile (
_extract_peaks_soundfile) _downsample_peaksshared helper produces list of floats 0.0-1.0lightsync/api/audio.pyhasGET /waveformendpoint- Waveform endpoint returns
{"peaks": [...], "count": N, "file": "..."} - Waveform extraction uses
asyncio.to_thread(non-blocking) - Returns 400 if no file loaded, 503 if engine not ready
peaksquery param clamped between 100-5000 </acceptance_criteria> Waveform extraction works for MP3/WAV/FLAC/OGG; HTTP endpoint returns peak data
Replace the transport-bar footer content:
<footer class="transport-bar" id="transport-panel">
<div class="transport-controls">
<button id="btn-play" class="transport-btn" title="Play">▶</button>
<button id="btn-pause" class="transport-btn" title="Pause" style="display:none">▮▮</button>
<button id="btn-stop" class="transport-btn" title="Stop">■</button>
</div>
<div class="transport-time">
<span id="time-current">0:00.0</span>
<span class="text-dim">/</span>
<span id="time-duration">0:00.0</span>
</div>
<div class="transport-seek">
<input type="range" id="seek-bar" min="0" max="100" value="0" step="0.1" class="seek-slider">
</div>
<div class="transport-file">
<input type="text" id="audio-path" placeholder="Audio file path on server..." class="audio-path-input">
<button id="btn-load" class="transport-btn" title="Load audio">LOAD</button>
</div>
</footer>
Replace the timeline-panel content (keep the panel div, replace inner content):
<main class="main-area">
<div class="panel" id="timeline-panel" style="flex: 1; display: flex; flex-direction: column;">
<div class="panel-header">TIMELINE</div>
<div class="waveform-container" style="flex: 1; position: relative; min-height: 100px;">
<canvas id="waveform-canvas" style="width: 100%; height: 100%;"></canvas>
<div id="playback-cursor" class="playback-cursor"></div>
</div>
</div>
</main>
- Update
lightsync/frontend/style.css— add transport and waveform styles:
Append these styles (preserve ALL existing styles):
/* Transport bar */
.transport-controls {
display: flex;
gap: 4px;
align-items: center;
}
.transport-btn {
background: var(--surface);
color: var(--accent);
border: 1px solid var(--border-bright);
padding: 4px 10px;
font-family: 'JetBrains Mono', monospace;
font-size: 0.85rem;
cursor: pointer;
min-width: 36px;
text-align: center;
}
.transport-btn:hover {
background: var(--accent);
color: var(--bg);
}
.transport-time {
font-family: 'JetBrains Mono', monospace;
font-size: 0.85rem;
white-space: nowrap;
min-width: 120px;
}
.transport-seek {
flex: 1;
display: flex;
align-items: center;
}
.seek-slider {
width: 100%;
accent-color: var(--accent);
cursor: pointer;
}
.transport-file {
display: flex;
gap: 4px;
align-items: center;
}
.audio-path-input {
background: var(--bg);
color: var(--text);
border: 1px solid var(--border);
padding: 4px 8px;
font-family: 'JetBrains Mono', monospace;
font-size: 0.8rem;
width: 200px;
}
.audio-path-input:focus {
border-color: var(--accent);
outline: none;
}
/* Transport bar layout */
.transport-bar {
display: flex;
gap: 12px;
align-items: center;
padding: 6px 12px;
}
/* Waveform */
.waveform-container {
background: var(--bg);
border: 1px solid var(--border);
overflow: hidden;
}
.playback-cursor {
position: absolute;
top: 0;
left: 0;
width: 2px;
height: 100%;
background: var(--accent);
pointer-events: none;
transition: left 0.1s linear;
z-index: 10;
}
IMPORTANT: Read style.css first to identify the existing CSS variable names. Use var(--accent), var(--bg), var(--surface), var(--text), var(--border), var(--border-bright), var(--text-dim) — whatever variable names are already defined. Do NOT create new color variables. The terminal/hacker aesthetic (UI-02) must be maintained.
- Update
lightsync/frontend/app.js— add transport logic, waveform rendering, and position tick handling:
Add AFTER the existing LightSyncClient class definition but BEFORE const client = new LightSyncClient():
// --- Waveform rendering ---
let waveformPeaks = [];
let audioDuration = 0;
let audioLoaded = false;
function drawWaveform(canvas, peaks) {
const ctx = canvas.getContext('2d');
const dpr = window.devicePixelRatio || 1;
const rect = canvas.getBoundingClientRect();
canvas.width = rect.width * dpr;
canvas.height = rect.height * dpr;
ctx.scale(dpr, dpr);
const w = rect.width;
const h = rect.height;
const mid = h / 2;
const barWidth = w / peaks.length;
ctx.clearRect(0, 0, w, h);
// Draw waveform bars
ctx.fillStyle = getComputedStyle(document.documentElement).getPropertyValue('--accent').trim() || '#0f0';
for (let i = 0; i < peaks.length; i++) {
const amp = peaks[i] * mid * 0.9;
const x = i * barWidth;
ctx.fillRect(x, mid - amp, Math.max(barWidth - 0.5, 0.5), amp * 2);
}
}
async function loadWaveform() {
try {
const res = await fetch('/api/audio/waveform?peaks=2000');
if (!res.ok) return;
const data = await res.json();
waveformPeaks = data.peaks;
const canvas = document.getElementById('waveform-canvas');
if (canvas) drawWaveform(canvas, waveformPeaks);
} catch (err) {
console.error('[waveform]', err);
}
}
// --- Time formatting ---
function formatTime(seconds) {
if (seconds == null || isNaN(seconds)) return '0:00.0';
const m = Math.floor(seconds / 60);
const s = seconds % 60;
return `${m}:${s < 10 ? '0' : ''}${s.toFixed(1)}`;
}
// --- Position update from WebSocket tick ---
function updatePosition(position, duration, paused) {
audioDuration = duration || 0;
document.getElementById('time-current').textContent = formatTime(position);
document.getElementById('time-duration').textContent = formatTime(duration);
// Update seek bar
const seekBar = document.getElementById('seek-bar');
if (duration > 0 && !seekBar._dragging) {
seekBar.value = (position / duration) * 100;
}
// Update playback cursor position
const cursor = document.getElementById('playback-cursor');
const container = document.querySelector('.waveform-container');
if (cursor && container && duration > 0) {
const pct = (position / duration) * 100;
cursor.style.left = pct + '%';
}
// Toggle play/pause button visibility
const btnPlay = document.getElementById('btn-play');
const btnPause = document.getElementById('btn-pause');
if (btnPlay && btnPause) {
btnPlay.style.display = paused ? '' : 'none';
btnPause.style.display = paused ? 'none' : '';
}
}
Modify the handleMessage method of LightSyncClient:
handleMessage(msg) {
if (msg.type === 'tick') {
updatePosition(msg.position, msg.duration, msg.paused);
audioLoaded = msg.loaded;
} else {
console.debug('[ws]', msg);
}
}
Add transport button event listeners AFTER client.connect() and loadDevices():
// --- Transport controls ---
document.getElementById('btn-play')?.addEventListener('click', () => {
client.send({ type: 'play' });
});
document.getElementById('btn-pause')?.addEventListener('click', () => {
client.send({ type: 'pause' });
});
document.getElementById('btn-stop')?.addEventListener('click', () => {
client.send({ type: 'pause' });
client.send({ type: 'seek', position: 0 });
});
// Seek bar interaction
const seekBar = document.getElementById('seek-bar');
if (seekBar) {
seekBar.addEventListener('mousedown', () => { seekBar._dragging = true; });
seekBar.addEventListener('mouseup', () => {
seekBar._dragging = false;
const pos = (parseFloat(seekBar.value) / 100) * audioDuration;
client.send({ type: 'seek', position: pos });
});
}
// Load audio button
document.getElementById('btn-load')?.addEventListener('click', async () => {
const pathInput = document.getElementById('audio-path');
const path = pathInput?.value.trim();
if (!path) return;
try {
const res = await fetch('/api/audio/load', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ path }),
});
if (!res.ok) {
const err = await res.json();
console.error('[audio] load failed:', err.detail);
return;
}
// Wait briefly for mpv to initialize, then fetch waveform
setTimeout(loadWaveform, 500);
} catch (err) {
console.error('[audio] load error:', err);
}
});
// Click on waveform to seek
document.querySelector('.waveform-container')?.addEventListener('click', (e) => {
if (audioDuration <= 0) return;
const rect = e.currentTarget.getBoundingClientRect();
const pct = (e.clientX - rect.left) / rect.width;
const pos = pct * audioDuration;
client.send({ type: 'seek', position: pos });
});
// Resize waveform on window resize
window.addEventListener('resize', () => {
if (waveformPeaks.length > 0) {
const canvas = document.getElementById('waveform-canvas');
if (canvas) drawWaveform(canvas, waveformPeaks);
}
});
IMPORTANT: Keep ALL existing code in app.js (LightSyncClient class, loadDevices, removeDevice, escapeHtml, device form handler). Only ADD the new transport/waveform code and modify handleMessage. cd /home/claude/led2 && grep -q "waveform-canvas" lightsync/frontend/index.html && grep -q "btn-play" lightsync/frontend/index.html && grep -q "playback-cursor" lightsync/frontend/index.html && grep -q "drawWaveform" lightsync/frontend/app.js && grep -q "updatePosition" lightsync/frontend/app.js && grep -q "formatTime" lightsync/frontend/app.js && grep -q "loadWaveform" lightsync/frontend/app.js && grep -q 'msg.type.*tick' lightsync/frontend/app.js && grep -q "playback-cursor" lightsync/frontend/style.css && grep -q "transport-btn" lightsync/frontend/style.css && echo "PASS" <acceptance_criteria>
index.htmltransport bar contains: play button (#btn-play), pause button (#btn-pause), stop button (#btn-stop), time display (#time-current, #time-duration), seek slider (#seek-bar), audio path input (#audio-path), load button (#btn-load)index.htmltimeline panel contains: waveform canvas (#waveform-canvas), playback cursor (#playback-cursor)app.jshasdrawWaveform(canvas, peaks)function that renders bars on canvas using --accent colorapp.jshasupdatePosition(position, duration, paused)that updates time display, seek bar, cursor, and play/pause button visibilityapp.jshandleMessage dispatchesmsg.type === 'tick'toupdatePositionapp.jsplay/pause/stop buttons send WebSocket commands viaclient.send()app.jsseek bar mouseup sends{type: "seek", position: N}with position calculated from bar value and durationapp.jsload button POSTs to/api/audio/loadthen callsloadWaveform()after 500ms delayapp.jsclick on waveform-container calculates position from click x-coordinate and sends seek commandstyle.csshas styles for.transport-btn,.transport-time,.seek-slider,.waveform-container,.playback-cursor- CSS uses existing variable names (--accent, --bg, --surface, --text, --border, --border-bright) — no new color variables
- All existing app.js code (device list, device form, LightSyncClient class) preserved
- Terminal/hacker aesthetic maintained (monospace fonts, dark theme) </acceptance_criteria> Transport controls work in browser; waveform renders from audio data; position cursor moves in real time
cd /home/claude/led2 && ~/bin/deploy.sh /home/claude/led2
Verification steps:
- Open https://lightsync.groll.cloud in browser (Authelia login if prompted)
- Place a test audio file on the server (e.g., copy to /app/ inside the container)
- Enter the file path in the audio path input box and click LOAD
- Verify: waveform appears in the timeline panel
- Click Play — verify position cursor moves, time display counts up
- Click Pause — verify playback stops, cursor stops
- Drag seek bar — verify cursor jumps to new position
- Click on waveform — verify cursor seeks to clicked position
- Let it play for 30+ seconds — verify no drift in position updates (ticks arrive smoothly)
- Check browser console for errors
Resume signal: Type "approved" or describe issues. cd /home/claude/led2 && grep -q "waveform-canvas" lightsync/frontend/index.html && grep -q "btn-play" lightsync/frontend/index.html && echo "PASS" User confirms: audio loads, waveform renders, play/pause/seek work, position cursor tracks playback in real time
- Waveform extraction works for WAV/FLAC/OGG (soundfile) and MP3 (ffmpeg pipe) - GET /api/audio/waveform returns peak data array - Transport bar has play/pause/stop/seek controls - Waveform canvas renders peaks with terminal-aesthetic styling - Position cursor moves at 10Hz matching WebSocket ticks - Click-to-seek works on waveform - All existing Phase 1 UI (devices panel, device form) still works<success_criteria>
- Waveform rendered from audio samples in timeline UI (AUD-05)
- Transport controls connected to WebSocket commands
- Position cursor tracks audio position in real time
- Terminal/hacker aesthetic maintained throughout
- Human verification confirms end-to-end audio playback works </success_criteria>