Files
yoloyolo/.planning/phases/02-audio-synthesis-engine/02-RESEARCH.md
T
2026-03-26 11:29:55 +01:00

680 lines
35 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Phase 2: Audio Synthesis Engine - Research
**Researched:** 2026-03-26
**Domain:** Go additive audio synthesis, stereo MP3 encoding, EMA amplitude smoothing, WAV intermediate format
**Confidence:** HIGH — decisions are locked, stack is pre-determined in CLAUDE.md, all API signatures verified against official pkg.go.dev docs.
---
<user_constraints>
## User Constraints (from CONTEXT.md)
### Locked Decisions
**Tone mapping**
- D-01: Wide ambient frequency range: 60-800 Hz across 11 traffic classes
- D-02: Frequencies use musical intervals (harmonic relationships — fifths, octaves, etc.) so layers blend pleasantly
- D-03: Approximate register assignments: ICMP ~65 Hz (deep bass), DNS ~110 Hz, HTTPS ~175 Hz, HTTP ~220 Hz, SSH ~330 Hz, SMTP ~440 Hz, NTP ~520 Hz, DHCP ~600 Hz, other-TCP/other-UDP ~700-800 Hz
- D-04: "unknown" traffic gets a dissonant/detuned tone — slightly off-key or beating frequency
**Waveform texture**
- D-05: Each drone layer uses fundamental + 2-3 harmonics (not pure sine waves)
- D-06: Harmonic ratios vary per traffic class for subtle timbral distinction
**Amplitude dynamics**
- D-07: EMA smoothing with medium responsiveness (1-2 second attack/decay)
- D-08: Layers never fade to full silence — once seen, stays at a fixed whisper floor (~2-5% of max amplitude)
- D-09: Fixed whisper floor (not recency-based decay)
**Mixing and output**
- D-10: Fixed equal gain per layer — each of 11 layers gets 1/11 of headroom (~0.09 max amplitude)
- D-11: Stereo output with register-based panning: bass center, mid spread L/R, high frequencies wider
- D-12: Panning positions: ICMP center, DNS slight-L, HTTPS slight-R, HTTP center-L, SSH center-R, SMTP mid-L, NTP mid-R, DHCP wide-L, other-TCP wide-R, other-UDP wide-L, unknown center
- D-13: 44100 Hz sample rate (CD quality)
- D-14: 128 kbps MP3 encoding via `sjzar/go-lame`
- D-15: `-o` flag for output path; defaults to `netsynth-<timestamp>.mp3` when omitted
- D-16: Empty (zero-packet) input produces a clear error message, not a corrupt or zero-byte MP3
### Claude's Discretion
- Exact Hz values per class (within the 60-800 Hz range, using musical intervals)
- Specific harmonic ratios per traffic class (within the "2-3 harmonics" constraint)
- Exact EMA alpha/decay constants to achieve the 1-2 second feel
- Exact whisper floor percentage (within 2-5% range)
- WAV intermediate format usage (per CLAUDE.md recommendation of go-audio/wav)
- Stereo panning implementation (constant-power pan law vs linear)
### Deferred Ideas (OUT OF SCOPE)
None — discussion stayed within phase scope.
</user_constraints>
---
<phase_requirements>
## Phase Requirements
| ID | Description | Research Support |
|----|-------------|------------------|
| SYNTH-01 | Each traffic class generates a distinct ambient/drone layer (layered sine/harmonic waves) | Additive synthesis pattern: oscillator struct with phase accumulator, harmonic table, per-class config |
| SYNTH-02 | Drone layer amplitudes evolve over time windows based on traffic volume per class | EMA update pattern: α derived from window duration; amplitude target driven by WindowSnapshot.Counts |
| SYNTH-03 | Multiple drone layers are mixed into a single coherent audio stream without distortion | Fixed-gain mixer: 1/11 per layer guarantees sum ≤ 1.0; constant-power stereo panning; go-audio/wav → go-lame pipeline |
| OUT-01 | User can specify output file path via `-o` flag (defaults to `netsynth-<timestamp>.mp3`) | Cobra `StringVarP` flag; `time.Now().Format` for timestamp default |
| OUT-02 | Output is encoded as a valid MP3 file | `sjzar/go-lame` LameWriter: set params → InitParams() → Write([]byte PCM) → Close() |
| OUT-03 | Empty captures (zero packets) produce a clear error instead of a corrupt file | Check `totalPackets == 0` before encoding; return descriptive `fmt.Errorf` instead |
</phase_requirements>
---
## Summary
Phase 2 builds the complete audio pipeline: oscillator bank → EMA amplitude smoother → stereo mixer → WAV buffer → LAME MP3 encoder. It operates against synthetic `classify.WindowSnapshot` inputs — no live capture is involved. All major library choices are locked in CLAUDE.md and prior research: `sjzar/go-lame` v0.0.9 for MP3, `go-audio/wav` for the intermediate format, hand-rolled additive synthesis, and Cobra for the `-o` flag.
The synthesis architecture is well-understood from Phase 1 research. The core math is: for each window, compute a target amplitude per traffic class from packet counts, advance the EMA toward that target, render N PCM samples (fundamental + 2-3 harmonics with class-specific ratios), apply constant-power stereo panning into L/R channels, mix the 11 layers, write stereo interleaved int16 samples to a WAV buffer, and encode to MP3 via go-lame at 44100 Hz / 128 kbps / 2 channels.
Two environmental blockers need Wave 0 tasks: `gcc` and `ffprobe` are not installed. `gcc` is required to build `sjzar/go-lame` (CGo). `ffprobe` is needed for the success criterion that validates the MP3 output. Both are available via `apt` on this machine.
**Primary recommendation:** Build `synth/`, `encode/`, and extend `cmd/` as three independent Go packages. Test each in isolation using table-driven tests with synthetic `WindowSnapshot` inputs and PCM buffer assertions before wiring to the full pipeline in Phase 3.
---
## Standard Stack
### Core (locked by CLAUDE.md)
| Library | Version | Purpose | Why Standard |
|---------|---------|---------|--------------|
| `github.com/sjzar/go-lame` | v0.0.9 | MP3 encoding via embedded LAME C source | Locked. Embeds libmp3lame; no system dep. CGO_ENABLED=1 required. |
| `github.com/go-audio/wav` | latest | WAV file I/O as PCM intermediate buffer | Locked. Simplifies PCM → LAME pipeline; decouples synthesis from encoding. |
| `github.com/spf13/cobra` | v1.10.2 | `-o` flag; already in go.mod | Already wired in main.go. |
| Hand-rolled additive synthesis | — | Oscillator, amplitude EMA, mixer | Locked. `go-audio/generator` archived 2026-02-01; no library adds value. |
**go-lame is NOT yet in go.mod.** Wave 0 must add it.
**go-audio/wav is NOT yet in go.mod.** Wave 0 must add it.
**Installation:**
```bash
# Requires PATH to Go binary (see environment notes)
go get github.com/sjzar/go-lame@v0.0.9
go get github.com/go-audio/wav@latest
```
---
## Architecture Patterns
### Recommended Package Structure
```
synth/
├── config.go # FreqConfig struct: frequency table, harmonic ratios, pan positions
├── oscillator.go # Phase-accumulator sine oscillator, Harmonics() render method
├── layer.go # Layer struct: Oscillator + EMA amplitude state per TrafficClass
├── bank.go # OscillatorBank: 11 layers, Update(WindowSnapshot), Render(N frames) → [][2]float64
└── mixer.go # Sum layers L+R, convert float64 → int16 interleaved (stereo)
encode/
└── mp3.go # NewMP3Encoder(path, sampleRate, bitrate) → writes WAV buffer → LAME encode → file
cmd/netsynth/
└── main.go # Add -o flag, zero-packet guard, wire synth bank → encode
```
### Pattern 1: Phase-Accumulator Oscillator with Harmonics
**What:** Each oscillator maintains a `phase float64` that advances by `freq/sampleRate` per sample, wrapping at 1.0. `math.Sin(2π * phase)` yields the fundamental. Harmonics at integer multiples use the same phase scaled by harmonic ratio.
**Why:** Avoids `math.Sin(2π * freq * t / sampleRate)` which loses floating-point precision over long runs. Phase accumulator stays near zero, preserving accuracy indefinitely.
**Example:**
```go
// Source: standard DSP practice; verified against Dylan Meeus audio-from-scratch series
type Oscillator struct {
phase float64
freq float64
sr float64
}
// Advance returns one sample (fundamental + harmonics summed and normalized).
func (o *Oscillator) Advance(harmonics []HarmonicDef) float64 {
sum := 0.0
totalWeight := 0.0
for _, h := range harmonics {
sum += h.Amplitude * math.Sin(2*math.Pi*o.phase*float64(h.Ratio))
totalWeight += h.Amplitude
}
o.phase += o.freq / o.sr
if o.phase >= 1.0 {
o.phase -= 1.0
}
if totalWeight > 0 {
return sum / totalWeight // normalize to [-1, 1]
}
return 0
}
```
### Pattern 2: EMA Amplitude Smoothing
**What:** On each `WindowSnapshot`, compute a `targetAmplitude` for each layer. The current amplitude moves toward target each audio sample using coefficient `α`. α is derived from desired time constant τ (in seconds) and sample rate SR: `α = 1 - exp(-1 / (τ * SR))`.
**Why 1-2 second feel:** With τ = 1.0s and SR = 44100, α ≈ 0.0000227. After τ seconds of samples, amplitude reaches ~63% of target — perceptually "medium responsive." For attack/decay symmetry, use the same α for both directions.
**Formula:**
```go
// Source: standard signal processing; EMA with continuous-time derivation
tau := 1.0 // seconds — Claude's discretion; adjust for 1-2s feel
alpha := 1.0 - math.Exp(-1.0/(tau*float64(sampleRate)))
// Per-sample update (inside render loop):
layer.currentAmp += alpha * (layer.targetAmp - layer.currentAmp)
```
**Whisper floor:** After receiving the first snapshot showing count > 0 for a class, set `layer.whisperFloor = 0.03` (3% — within the 2-5% range). The target amplitude is:
```go
target := whisperFloor + (1.0 - whisperFloor) * normalizedRate
// where normalizedRate = float64(count) / float64(maxCountSeen)
```
### Pattern 3: Constant-Power Stereo Panning
**What:** Given a pan position `p ∈ [-1, 1]` where -1 = full left, 0 = center, +1 = full right, the per-channel gains are:
```go
// Source: standard audio engineering — constant-power (equal-power) pan law
angle := (p + 1.0) / 2.0 * math.Pi / 2.0 // map [-1,1] → [0, π/2]
gainL := math.Cos(angle)
gainR := math.Sin(angle)
```
Constant-power panning preserves perceived loudness as the signal moves across the stereo field. Linear panning creates a loudness dip at center. Given the ambient nature of this output, constant-power is correct.
**Pan table (from D-12):**
```go
var PanPositions = map[classify.TrafficClass]float64{
classify.ClassICMP: 0.0, // center
classify.ClassDNS: -0.2, // slight-L
classify.ClassHTTPS: +0.2, // slight-R
classify.ClassHTTP: -0.35, // center-L
classify.ClassSSH: +0.35, // center-R
classify.ClassSMTP: -0.55, // mid-L
classify.ClassNTP: +0.55, // mid-R
classify.ClassDHCP: -0.75, // wide-L
classify.ClassOtherTCP: +0.75, // wide-R
classify.ClassOtherUDP: -0.75, // wide-L (matches DHCP register)
classify.ClassUnknown: 0.0, // center (stand out via dissonance, not position)
}
```
### Pattern 4: WAV Buffer → LAME Pipeline
**What:** Render all PCM samples to an `audio.IntBuffer` (via `go-audio/wav`), write to a `bytes.Buffer` as WAV, then pass the raw PCM bytes to `sjzar/go-lame`'s `LameWriter.Write()`.
**Why WAV intermediate:** Decouples synthesis timing from LAME's CGo overhead. The entire PCM block is available before encoding starts, which is correct for a batch-output tool. WAV also serves as a debug artifact — save it alongside the MP3 during development.
**go-lame API (verified against pkg.go.dev):**
```go
// Source: pkg.go.dev/github.com/sjzar/go-lame — verified 2026-03-26
import "github.com/sjzar/go-lame"
outFile, _ := os.Create(outputPath)
defer outFile.Close()
wr := lame.NewWriter(outFile)
wr.Encoder.SetInSamplerate(44100)
wr.Encoder.SetOutSamplerate(44100)
wr.Encoder.SetNumChannels(2) // stereo
wr.Encoder.SetBitrate(128)
wr.Encoder.SetQuality(5)
wr.Encoder.InitParams() // MUST call after all Set* calls
// PCM input format: int16 little-endian interleaved stereo bytes
// Render: []int16{L0, R0, L1, R1, ...} → convert to []byte → Write
pcmBytes := int16SliceToBytes(interleavedSamples)
wr.Write(pcmBytes)
wr.Close() // flushes LAME internal buffer, writes final MP3 frames
```
**CRITICAL:** `InitParams()` must be called after all configuration `Set*` calls and before the first `Write`. Skipping it produces corrupted MP3 output.
**int16 conversion:**
```go
import "encoding/binary"
func int16SliceToBytes(samples []int16) []byte {
buf := make([]byte, len(samples)*2)
for i, s := range samples {
binary.LittleEndian.PutUint16(buf[i*2:], uint16(s))
}
return buf
}
```
### Pattern 5: Harmonic Config Table
**What:** A static table maps each `TrafficClass` to its `FreqConfig`. Defined in `synth/config.go`. No magic numbers elsewhere.
**Recommended frequency assignments (Claude's discretion, within D-01 to D-04 constraints):**
The user specified approximate Hz values. Musical intervals starting from 65 Hz (ICMP as the bass root):
- 65 Hz is close to C2 (musical bass). Intervals: fifth = ×1.5, octave = ×2, major third ≈ ×1.25, minor third ≈ ×1.2.
```go
// Source: derived from D-03 constraints using musical interval ratios
type HarmonicDef struct {
Ratio int // harmonic number (1 = fundamental, 2 = octave, 3 = fifth+octave)
Amplitude float64 // relative weight
}
type FreqConfig struct {
BaseHz float64
Harmonics []HarmonicDef
Pan float64 // from PanPositions table
}
var ClassFreqConfigs = map[classify.TrafficClass]FreqConfig{
// Deep bass — pure, minimal harmonics (pad feel)
classify.ClassICMP: {65.0, []HarmonicDef{{1, 1.0}, {2, 0.4}, {3, 0.15}}, 0.0},
// Sub-bass — slightly brighter
classify.ClassDNS: {110.0, []HarmonicDef{{1, 1.0}, {2, 0.5}, {3, 0.25}}, -0.2},
// Warm mid-bass — HTTPS is the most common traffic, warm pad
classify.ClassHTTPS: {175.0, []HarmonicDef{{1, 1.0}, {2, 0.6}, {3, 0.3}}, 0.2},
// Mid — HTTP slightly brighter than HTTPS
classify.ClassHTTP: {220.0, []HarmonicDef{{1, 1.0}, {2, 0.5}, {4, 0.2}}, -0.35},
// SSH — buzzy (odd harmonics emphasized for character)
classify.ClassSSH: {330.0, []HarmonicDef{{1, 1.0}, {3, 0.6}, {5, 0.3}}, 0.35},
// SMTP — clean tone
classify.ClassSMTP: {440.0, []HarmonicDef{{1, 1.0}, {2, 0.3}, {3, 0.1}}, -0.55},
// NTP — high, pure
classify.ClassNTP: {520.0, []HarmonicDef{{1, 1.0}, {2, 0.25}}, 0.55},
// DHCP — wide stereo, brighter
classify.ClassDHCP: {600.0, []HarmonicDef{{1, 1.0}, {2, 0.35}, {3, 0.15}}, -0.75},
// other-TCP — high, minimal harmonics
classify.ClassOtherTCP: {700.0, []HarmonicDef{{1, 1.0}, {2, 0.2}}, 0.75},
// other-UDP
classify.ClassOtherUDP: {780.0, []HarmonicDef{{1, 1.0}, {2, 0.2}}, -0.75},
// unknown — detuned (beating frequency: fundamental + 1.03× = ~3Hz beat at 440Hz)
classify.ClassUnknown: {437.0, []HarmonicDef{{1, 1.0}, {2, 0.8}, {3, 0.4}}, 0.0},
// The 437 Hz combined with its 2nd harmonic at 874 Hz creates beating against
// any SMTP (440 Hz) present in the mix, reinforcing "something doesn't belong"
}
```
Note on "unknown" detuning: 437 Hz is ~12 cents flat from A4 (440 Hz). If SMTP traffic is present, the 3 Hz beat between 437 and 440 is perceptually jarring. The 2nd harmonic at 874 Hz also beats against the SMTP fundamental's second harmonic at 880 Hz — a compound dissonance. This is the correct implementation of D-04.
### Pattern 6: Per-Window Render Loop
**What:** For each received `WindowSnapshot`, update amplitude targets, then render exactly `samplesPerWindow` PCM frames (= `sampleRate * windowDurationMs / 1000`).
```go
// samplesPerWindow = 44100 * 500 / 1000 = 22050 samples per 500ms window
// Each frame = 2 int16 values (L, R) for stereo
samplesPerWindow := sampleRate * windowMs / 1000
func (b *OscillatorBank) RenderWindow(snap classify.WindowSnapshot) [][2]float64 {
b.updateTargets(snap) // sets targetAmp per layer from snap.Counts
frames := make([][2]float64, samplesPerWindow)
for i := range frames {
var sumL, sumR float64
for _, layer := range b.layers {
sample := layer.osc.Advance(layer.config.Harmonics)
layer.currentAmp += layer.alpha * (layer.targetAmp - layer.currentAmp)
gainL, gainR := panGains(layer.config.Pan)
sumL += sample * layer.currentAmp * layer.gainPerLayer * gainL
sumR += sample * layer.currentAmp * layer.gainPerLayer * gainR
}
frames[i] = [2]float64{sumL, sumR}
}
return frames
}
```
**gainPerLayer = 1.0 / 11 ≈ 0.0909** (D-10). With 11 layers each at max amplitude 1.0 and gain 1/11, the theoretical maximum sum is exactly 1.0. Clipping is impossible by construction — no dynamic limiter needed.
### Anti-Patterns to Avoid
- **Calling InitParams() before all Set* calls:** go-lame silently uses defaults for any parameters set after InitParams(). Always configure fully, then call InitParams() once.
- **Using float32 PCM samples directly with go-lame:** go-lame's `Encode` and `Write` expect `[]byte` representing int16 little-endian PCM. Convert float64 samples to int16 with `int16(sample * 32767)` and then to bytes via `encoding/binary`.
- **Writing an empty MP3:** If zero snapshots are received (D-16 / OUT-03), detect this before opening the output file. Opening the file before checking creates a zero-byte file on disk even if encoding fails.
- **Streaming MP3 mid-run:** MP3 VBR headers require finalization. Buffer all PCM then encode at the end. This is the correct pattern for non-interactive output (see REQUIREMENTS.md Out of Scope).
- **Phase accumulator overflow:** Use `phase -= 1.0` (not `math.Mod`) when phase exceeds 1.0. `math.Mod` is slower and introduces floating-point artifacts.
- **go-audio/wav for the final output file:** go-audio/wav requires an `io.WriteSeeker` (must support `Seek`). A `bytes.Buffer` does not implement `Seek`. Use `os.File` or `bytes.Buffer` + manual WAV header if only using WAV as intermediate. Recommended: render PCM to `[]int16`, skip WAV file entirely for production path, write bytes directly to LameWriter.
---
## Don't Hand-Roll
| Problem | Don't Build | Use Instead | Why |
|---------|-------------|-------------|-----|
| MP3 encoding | Custom MP3 framer | `sjzar/go-lame` v0.0.9 | MP3 frame format is complex; ID3 headers, bit reservoir, psychoacoustic model — thousands of edge cases |
| WAV file format | Custom WAV header writer | `go-audio/wav` (optional) or raw int16 bytes direct to go-lame | WAV header has mandatory chunk offsets that must be backpatched; go-audio/wav handles this |
| Musical interval math | Frequency ratio lookup tables | Use the ratios directly: fifth = ×1.5, octave = ×2 | Pre-computed in config table; no runtime math needed |
---
## Common Pitfalls
### Pitfall 1: InitParams() Not Called
**What goes wrong:** `LameWriter.Write()` silently produces empty or corrupted MP3 frames. No panic or error is returned. The MP3 file is created but fails `ffprobe` validation.
**Why it happens:** go-lame wraps libmp3lame's C API which requires explicit parameter initialization. The Go wrapper does not auto-call it.
**How to avoid:** In `encode/mp3.go`, make `NewMP3Encoder()` call `InitParams()` internally after all configuration. Callers never need to know about it.
**Warning signs:** MP3 file is non-zero bytes but `ffprobe` reports "Invalid data found when processing input" or duration is 0.
### Pitfall 2: Stereo Interleaving Order
**What goes wrong:** Left and right channels are swapped or mono is written when stereo is expected. Oscillators that should be wide-right sound wide-left.
**Why it happens:** go-lame expects `[L0, R0, L1, R1, ...]` interleaved. If samples are written as `[L0, L1, ..., R0, R1, ...]` (planar), the decoder interprets the first N/2 bytes as alternating L/R.
**How to avoid:** Interleave explicitly in the frame render loop: `pcm[i*2] = leftSample; pcm[i*2+1] = rightSample`.
### Pitfall 3: EMA Never Reaches Whisper Floor on First Window
**What goes wrong:** During the first window, `currentAmp` starts at 0. If a class has count > 0, the target is above the whisper floor but `currentAmp` starts from 0 — it takes several windows to ramp up. The first second of audio is unnaturally quiet.
**Why it happens:** EMA is a smoothing filter — it takes time to respond from zero.
**How to avoid:** Initialize `currentAmp = whisperFloor` for all layers at startup (not 0). This way all layers start at their floor, not silence. D-08 specifies the floor is always present once a class has been seen — pre-setting to whisper floor on construction is consistent with the intent.
### Pitfall 4: go-audio/wav Requires io.WriteSeeker
**What goes wrong:** Passing `bytes.Buffer` as the writer to `wav.NewEncoder` causes a compile error or runtime panic because `bytes.Buffer` does not implement `Seek`.
**Why it happens:** WAV format must backpatch header chunk sizes after the data is written. This requires seeking backwards in the file.
**How to avoid:** Either write PCM bytes directly to `LameWriter` (skip WAV entirely) or use a `bytes.Buffer` subtype that supports seeking. The simplest production path: render float64 samples → convert to int16 → convert to []byte → `LameWriter.Write()`. The WAV intermediate is optional for debugging.
### Pitfall 5: Zero-Packet Guard Must Precede File Creation
**What goes wrong:** The output file (`netsynth-<timestamp>.mp3`) is created on disk, then the zero-packet check fires and returns an error. A zero-byte file is left behind. Confusing to users.
**Why it happens:** Typical pattern opens the file first, then validates inputs.
**How to avoid:** Collect all `WindowSnapshot` values from the channel into a slice. Check `totalPackets == 0` across all snapshots. Only then open the output file and begin encoding. In Phase 2's isolated testing, this check is performed before passing snapshots to the encoder.
### Pitfall 6: gcc Not Installed
**What goes wrong:** `go build` with `CGO_ENABLED=1` fails with `cgo: C compiler "gcc" not found: exec: "gcc": executable file not found in $PATH`.
**Why it happens:** `sjzar/go-lame` embeds LAME C source and requires CGo compilation. The build environment on this machine does not have `gcc` installed.
**How to avoid:** Wave 0 task must install gcc: `sudo apt-get install -y gcc`.
---
## Environment Availability
| Dependency | Required By | Available | Version | Fallback |
|------------|------------|-----------|---------|----------|
| Go toolchain | All compilation | ✗ | — | Must install; available via apt or tarball |
| gcc / C compiler | `sjzar/go-lame` CGo build | ✗ | — | None — required for go-lame; install via apt |
| ffprobe | OUT-02 test validation (success criterion: "passes ffprobe validation") | ✗ | — | None for automated test; install via apt (ffmpeg package) |
| `sjzar/go-lame` v0.0.9 | MP3 encoding | ✗ (not in go.mod) | — | Wave 0: `go get github.com/sjzar/go-lame@v0.0.9` |
| `go-audio/wav` | WAV intermediate (optional) | ✗ (not in go.mod) | — | Can skip WAV and write PCM bytes directly to LameWriter |
**Note on Go toolchain:** The project's go.mod exists and was used to build Phase 1 successfully. STATE.md records `[Phase 01]: Go installed to /home/dev/tools/go-install/go (no sudo); PATH export required each session`. However, `/home/dev/tools/` does not currently exist on this machine. The Go binary path needs to be confirmed before Wave 0 executes. The module cache location will also affect where `go get` downloads packages.
**Missing dependencies with no fallback:**
- `gcc` — blocks go-lame CGo compilation. Install: `sudo apt-get install -y gcc`
- `ffprobe` — blocks MP3 validation in tests. Install: `sudo apt-get install -y ffmpeg`
- Go toolchain — path unknown; needs confirmation before Wave 0
**Missing dependencies with fallback:**
- `go-audio/wav` — can be skipped; write PCM bytes directly to LameWriter (saves one dependency)
---
## Validation Architecture
### Test Framework
| Property | Value |
|----------|-------|
| Framework | Go standard `testing` package (no external framework) |
| Config file | none — uses `go test ./...` |
| Quick run command | `go test ./synth/... ./encode/...` |
| Full suite command | `go test ./...` |
### Phase Requirements → Test Map
| Req ID | Behavior | Test Type | Automated Command | File Exists? |
|--------|----------|-----------|-------------------|-------------|
| SYNTH-01 | Each class produces a non-zero PCM output with distinct frequency content | unit | `go test ./synth/... -run TestOscillatorDistinct` | ❌ Wave 0 |
| SYNTH-01 | Harmonics produce richer waveform than pure sine | unit | `go test ./synth/... -run TestHarmonics` | ❌ Wave 0 |
| SYNTH-02 | EMA amplitude rises toward target over N samples | unit | `go test ./synth/... -run TestEMAAplitudeRise` | ❌ Wave 0 |
| SYNTH-02 | Whisper floor prevents amplitude reaching zero | unit | `go test ./synth/... -run TestWhisperFloor` | ❌ Wave 0 |
| SYNTH-03 | Mixer sum of 11 max-amplitude layers does not exceed 1.0 | unit | `go test ./synth/... -run TestMixerNoClip` | ❌ Wave 0 |
| SYNTH-03 | Stereo output has distinct L/R channel values for panned sources | unit | `go test ./synth/... -run TestStereoPan` | ❌ Wave 0 |
| OUT-01 | `-o flag` sets output path; default contains timestamp | unit | `go test ./cmd/... -run TestOutputFlag` | ❌ Wave 0 |
| OUT-02 | Encoded bytes form a valid MP3 (ffprobe check) | integration | `go test ./encode/... -run TestMP3Valid` | ❌ Wave 0 |
| OUT-03 | Zero-packet input returns error, no file created | unit | `go test ./encode/... -run TestZeroPacketError` | ❌ Wave 0 |
### Sampling Rate
- **Per task commit:** `go test ./synth/...`
- **Per wave merge:** `go test ./...`
- **Phase gate:** Full suite green before `/gsd:verify-work`
### Wave 0 Gaps
- [ ] `synth/config_test.go` — covers SYNTH-01 (oscillator distinct, harmonics)
- [ ] `synth/layer_test.go` — covers SYNTH-02 (EMA rise, whisper floor)
- [ ] `synth/mixer_test.go` — covers SYNTH-03 (no clip, stereo pan)
- [ ] `encode/mp3_test.go` — covers OUT-02, OUT-03
- [ ] Environment setup: `sudo apt-get install -y gcc ffmpeg` — required before any build
- [ ] `go get github.com/sjzar/go-lame@v0.0.9` — add to go.mod
- [ ] `go get github.com/go-audio/wav@latest` — add to go.mod (if using WAV intermediate)
---
## Code Examples
### Full go-lame Encode Path (Verified)
```go
// Source: pkg.go.dev/github.com/sjzar/go-lame — verified 2026-03-26
import (
"encoding/binary"
"os"
"github.com/sjzar/go-lame"
)
func EncodeMP3(outputPath string, stereoSamples [][2]float64, sampleRate int) error {
f, err := os.Create(outputPath)
if err != nil {
return err
}
defer f.Close()
wr := lame.NewWriter(f)
wr.Encoder.SetInSamplerate(sampleRate)
wr.Encoder.SetOutSamplerate(sampleRate)
wr.Encoder.SetNumChannels(2)
wr.Encoder.SetBitrate(128)
wr.Encoder.SetQuality(5)
wr.Encoder.InitParams() // MUST be called after all Set*
// Convert [][2]float64 → interleaved int16 bytes
pcm := make([]byte, len(stereoSamples)*4) // 2 channels * 2 bytes/sample
for i, frame := range stereoSamples {
l := int16(frame[0] * 32767)
r := int16(frame[1] * 32767)
binary.LittleEndian.PutUint16(pcm[i*4:], uint16(l))
binary.LittleEndian.PutUint16(pcm[i*4+2:], uint16(r))
}
if _, err := wr.Write(pcm); err != nil {
return err
}
return wr.Close()
}
```
### EMA Alpha Calculation
```go
// Source: standard continuous-time EMA derivation
// tau = time constant in seconds (1.0 for 1-second attack/decay feel)
// sampleRate = 44100
func emaAlpha(tau float64, sampleRate int) float64 {
return 1.0 - math.Exp(-1.0/(tau*float64(sampleRate)))
}
// Result: alpha ≈ 0.0000227 for tau=1.0, SR=44100
// After 44100 samples (1 second): amplitude reaches 63.2% of target
// After 3 seconds: 95% of target — consistent with "medium responsiveness" D-07
```
### Amplitude Target from WindowSnapshot
```go
// Source: derived from D-07, D-08, D-09
func (b *OscillatorBank) updateTargets(snap classify.WindowSnapshot) {
// Find max count across all classes for normalization
var maxCount int64
for _, count := range snap.Counts {
if count > maxCount {
maxCount = count
}
}
for _, class := range classify.AllClasses() {
layer := b.layers[class]
count := snap.Counts[class]
// Mark class as "seen" if it has appeared at all
if count > 0 {
layer.seen = true
}
var target float64
if layer.seen {
normalizedRate := 0.0
if maxCount > 0 {
normalizedRate = float64(count) / float64(maxCount)
}
target = layer.whisperFloor + (1.0-layer.whisperFloor)*normalizedRate
}
// If never seen: target stays 0.0 (no whisper floor yet)
layer.targetAmp = target
}
}
```
### Zero-Packet Guard (OUT-03)
```go
// Collect all snapshots before encoding
func RunSynthesis(snapshots <-chan classify.WindowSnapshot, outputPath string) error {
var allSnaps []classify.WindowSnapshot
var totalPackets int64
for snap := range snapshots {
allSnaps = append(allSnaps, snap)
totalPackets += snap.TotalPackets
}
if totalPackets == 0 {
return fmt.Errorf("no packets captured: output MP3 not written (empty capture produces no audio)")
}
// Now safe to open output file and encode
return encodeAllSnaps(allSnaps, outputPath)
}
```
### Output Flag Wiring in main.go
```go
// Add to main.go — following existing Cobra pattern
var outputPath string
rootCmd.Flags().StringVarP(&outputPath, "output", "o", "", "Output MP3 file path (default: netsynth-<timestamp>.mp3)")
// In run() function — resolve default before synthesis:
if outputPath == "" {
outputPath = fmt.Sprintf("netsynth-%s.mp3", time.Now().Format("20060102-150405"))
}
```
---
## State of the Art
| Old Approach | Current Approach | When Changed | Impact |
|--------------|------------------|--------------|--------|
| `go-audio/generator` for oscillators | Hand-rolled phase accumulator | Feb 2026 (archived) | No library dependency; ~20 lines replaces it |
| `google/gopacket` | `gopacket/gopacket` community fork | 2024 | Already using correct fork (in go.mod) |
| Dynamic-linked libmp3lame | `sjzar/go-lame` embedded C source | April 2025 (v0.0.9) | Self-contained CGo; no system package required |
**Deprecated/outdated:**
- `go-audio/generator`: Archived February 2026. Do not use. Already documented in CLAUDE.md.
- `faiface/beep`: Real-time playback focus — irrelevant for batch file output.
---
## Open Questions
1. **Go toolchain location on this machine**
- What we know: STATE.md says Go was installed to `/home/dev/tools/go-install/go` and PATH export is required each session. But `/home/dev/tools/` does not exist.
- What's unclear: Was Go installed in this container/environment differently, or does it need re-installation?
- Recommendation: Wave 0 first task should confirm `which go` or install Go fresh. The `go.mod` and `go.sum` exist, suggesting Go was successfully used for Phase 1.
2. **go-audio/wav vs raw PCM bytes to go-lame**
- What we know: CLAUDE.md recommends go-audio/wav as intermediate. The WAV library requires `io.WriteSeeker` which `bytes.Buffer` does not satisfy.
- What's unclear: Whether the planner should include go-audio/wav as a dependency or go direct PCM → LAME.
- Recommendation: Write PCM as `[]int16` directly to a `[]byte` buffer and pass to `LameWriter.Write()`. This is simpler, eliminates one dependency, and avoids the `io.WriteSeeker` complication. Reserve go-audio/wav for if a WAV debug output feature is wanted later.
3. **Subjective listening validation**
- What we know: STATE.md flags "Frequency mapping requires subjective listening validation — specific Hz values not determined by research; must test during Phase 2".
- What's unclear: The plan should allocate a manual listening step. Automated tests can verify PCM is non-zero and distinct across classes, but perceptual quality requires human ears.
- Recommendation: Include a Wave N task: "Manual listening test — play generated MP3 with synthetic traffic, verify each class is audibly distinct and the stereo field is perceived."
---
## Project Constraints (from CLAUDE.md)
The following directives from `CLAUDE.md` constrain this phase. Plans MUST NOT contradict them.
| Directive | Impact on Phase 2 |
|-----------|-------------------|
| Language: Go — single binary output | No new runtime dependencies; go-lame CGo is compile-time only |
| Audio format: MP3 output (not WAV or raw PCM) | WAV is intermediate only (optional); final output must be .mp3 |
| Non-interactive capture | All PCM is buffered then encoded after capture ends; no streaming |
| `sjzar/go-lame` v0.0.9 — use this, not viert/go-lame or shine-mp3 | Encoder package must import `github.com/sjzar/go-lame` |
| `go-audio/generator` is archived — do NOT use | Oscillator must be hand-rolled (phase accumulator pattern) |
| `faiface/beep`, `dasa.cc/snd` — do NOT use | No real-time audio libraries |
| `go-audio/wav` recommended for intermediate format | Use if WAV intermediate is included; or skip entirely (see Open Question 2) |
| CGO_ENABLED=1 required for go-lame build | Build commands must set CGO_ENABLED=1; gcc must be installed |
| GSD workflow enforcement — use `/gsd:execute-phase` | No direct file edits outside GSD workflow |
---
## Sources
### Primary (HIGH confidence)
- `pkg.go.dev/github.com/sjzar/go-lame` — API signatures, InitParams requirement, LameWriter pattern, PCM byte format — verified 2026-03-26
- `pkg.go.dev/github.com/go-audio/wav` — Encoder API, io.WriteSeeker requirement — verified 2026-03-26
- `/home/dev/workspace/yoloyolo/classify/types.go` — TrafficClass constants, WindowSnapshot struct — direct codebase read
- `/home/dev/workspace/yoloyolo/aggregate/window.go` — Aggregate() signature, channel contract — direct codebase read
- `/home/dev/workspace/yoloyolo/cmd/netsynth/main.go` — existing Cobra flag pattern, pipeline wiring — direct codebase read
- `/home/dev/workspace/yoloyolo/CLAUDE.md` — locked stack decisions, forbidden libraries — direct codebase read
### Secondary (MEDIUM confidence)
- `.planning/research/ARCHITECTURE.md` — pipeline patterns, per-layer amplitude lerp, anti-patterns — project research document from 2026-03-24
- `.planning/research/STACK.md` — library rationale, audio synthesis architecture note — project research document from 2026-03-24
- Standard DSP literature — EMA formula `α = 1 - exp(-1/(τ·SR))`, constant-power pan law `cos/sin` — mathematically stable, textbook-level
### Tertiary (LOW confidence)
- Frequency assignments (specific Hz values) — derived from D-03 constraints using musical interval ratios; require subjective listening validation
---
## Metadata
**Confidence breakdown:**
- Standard stack: HIGH — locked in CLAUDE.md, APIs verified against pkg.go.dev
- Architecture: HIGH — patterns verified from Phase 1 codebase; math is standard DSP
- Pitfalls: HIGH — InitParams trap verified against go-lame API; others from direct API inspection
- Frequency assignments: LOW — require subjective listening test; values are reasonable starting points
**Research date:** 2026-03-26
**Valid until:** 2026-06-26 (stable stack; go-lame and go-audio/wav APIs are stable)