Commit b747cc70 by PLN (Algolia)

merge: the CosmicFest release pipeline, and a shelf that cannot collide

PLN: "lets merge on main all these tooling improvements and go on".

17 commits, 16 files, all reviewed as a diff against master before merging and
not just at the tip — the lesson from the 200-file merge in ea52a32c..2f947c2c.
Twelve are new files, four are edits: build_release_plan.py, catalog_view.json
(a regeneration, 73 -> 81 tracks, dropping none), the CosmicFest spec, and
TASKS_DUMP.md.

What it adds, in one line each:

  master_stemless.py     a bus-only master for a gig with no stems
  demucs_sections.py     separation in bounded chunks, sections and global
  eda_stems.py           the required stem EDA, from separations instead
  setlist_to_segments.py an ear-called setlist becomes segments
  compare_masters.py     an A/B in numbers, so the ear spends itself on taste
  build_gig_tracksjson.py  a gig page generated from sourced facts
  check_permalinks.py    the shelf is one namespace; refuse a collision
  finish_cosmic.sh       everything downstream of the masters, unattended

Delivered: two masters in spec (-14.0 LUFS, -0.9 dBTP, 44.1/24, duration exact
to a microsecond against the source), a 14-track split that verifies clean, 14
section stem sets plus a global pass, and three EDA questions answered — the
missing air is cymbals, the kick is present 44.8% of the set against OPAL's
59%, and the global separation pass is a measured waste of time (corr
0.943-0.977, level deltas under a quarter of a dB).

Four things I got wrong are recorded in the branch rather than tidied out of it,
because three shared one shape worth naming: a check whose failure looks exactly
like success. A dropped rtk grep plus a block-buffered log made a live render
look dead, and two processes wrote the same paths for forty minutes. A speed
claim counted section directories, which appear when a writer OPENS. An edit
no-op'd while printing that it had worked. The fourth was not new — the first
full render came out at 192 kHz, a loudnorm trap this project had already
documented and I did not read before writing a new chain.

Not merged and still PLN's: the +5 vs +8 air-shelf call, and a release_signoff.
The CosmicSet upload on the shelf right now is PRIVATE and its plan says
_UNSIGNED.
parents 2f947c2c 34b02925
......@@ -1128,3 +1128,190 @@ words where they exist. IDs are stable; reference them in commits.
paint-reset culprit; that mystery is REOPENED with no suspect.
- Midi-Thru→SC re-add explained: SuperCollider's own `MIDIIn.connectAll`.
- A1 label off-by-one DISPROVEN ("d9 level [ard]" is correct).
---
## AMENDMENT — 2026-09-01, the two-set release board *(session-planning)*
PLN's declared goal, verbatim: *"release the tracks finish both Opal set and Cosmic
set release of both 1. their one live club mix, and 2. the album split"* — plus an
acknowledgement that sets the CosmicFest bar: *"i ack cosmicset is a raw record from
stage with effects, we can only do our best, 80/20 likely"*.
So the north star is now **four deliverables, not one**: {OPAL, COSMIC} x {continuous
club mix, per-track album split}.
### The scorecard (measured this session, not recalled)
| | OPAL-26 "Sunset Forest" | COSMICFEST-26 |
|---|---|---|
| source | 12 orbit stems + master, 48 k/24-bit | **ONE stage mic MP3**, 320 k/44.1, 62:59 |
| source location | `Prod/Opal26_master/` | **`~/Downloads/ZOOM0067.MP3` — single copy** |
| boundaries | ✅ ear-signed 2026-08-16 | ✅ ear pass 2026-08-29 (**10** decided + 8 skipped) |
| setlist | 14 tracks (Desire cut) | 14 tracks, `cosmicfest-2026_setlist_ear.json` |
| gig spec | ✅ `judge_specs/opal26.json` | 🔴 **DOES NOT EXIST** |
| segments | ✅ `segments_v4.json` | 🔴 none |
| mix / master | ✅ v4 premix, streaming + club | 🔴 nothing rendered |
| album split | ✅ 14 x 2 variants | 🔴 none |
| club mix | ✅ 2 x `*_nodrop.flac`, 72.7 min | 🔴 none |
| ear review of the MIX | ✅ A1.7, on the rendered files | ⛔ impossible — no mix to review |
| published | 🟡 8/15 SC, private, 2 defects | 🔴 nothing, and no www page |
**Answering the four questions directly:** boundaries are DONE for *both* sets. The
mix is done and ear-reviewed for OPAL *only*. CosmicFest has had zero mastering work —
its ear pass covered *where the cuts are*, never *how it sounds*.
### ✅ R-0 DONE 2026-09-01 · the CosmicFest master is now backed up
Was `~/Downloads/ZOOM0067.MP3`, 144 MB, a single copy of an irreplaceable gig.
Now two verified copies, `sha256 7a197589…d671b` identical on both, Downloads
original removed only after both were confirmed:
/home/pln/Work/Sound/Prod/Cosmic26_master/ZOOM0067.MP3
/mnt/freebox/PLN/Work/Sound/Prod/Cosmic26_master/ZOOM0067.MP3
(Note the naming trap: `Prod/cosmicfest/` and `Prod/cosmicfestv0.live1.*` are the
**2025** edition — files dated 2025-06-27/28. Nothing there belongs to this release.)
### 📏 The CosmicFest source is far better than "raw stage mic" implies
Measured with `ebur128` on the full 62:59: **I = −25.2 LUFS · true peak −5.1 dBFS ·
LRA 11.1 LU**. Read that as three pieces of good news:
- **−5.1 dBFS peak** = it was recorded conservatively and is **not clipped**. There is
real headroom to work with, not a squashed brick.
- **LRA 11.1 LU** = the FOH chain did **not** crush the dynamics. A mangled desk feed
typically lands under 6 LU.
- **−25.2 LUFS** = it is simply *quiet*. Reaching −14 LUFS streaming needs ~+11 dB,
which will demand limiting, but that is a normal gain problem, not a repair job.
The 80/20 is therefore genuinely reachable: gain-stage, broadband/room EQ,
harshness control, limit. **No premix is possible** — one stereo file, no stems, so
none of OPAL's per-orbit surgery (the d4 trim, the d5 cut) has an analogue here.
### ✅ KEYSTONE C-0 DONE 2026-09-01 — `judge_specs/cosmicfest-2026.json` exists
**and the stemless path is VERIFIED, not assumed**
`apply_boundaries` ran clean on it: 14 tracks, 63.0 min of a 63.0 min master, no
gaps, every assertion passed, `segments_v1.json` written. Confirmed by reading the
code rather than by trying: `apply_boundaries` and `render_release` touch **no** stem
key at all — only `segments`, `master`, `variants`, `releaseRoot`, `releaseTag`. The
one tool that genuinely needs stems is `build_judge_set`, which is precisely why the
nominal segments came from the new `setlist_to_segments.py` instead.
That converter is the reusable part: it turns a `_setlist_ear.json` into nominal
segments plus an `apply_boundaries`-schema ear file, and it keeps the two grades of
truth apart so a playhead call and an inferred edge can never be confused later.
Titles resolve authored → OPAL-release → catalog → `NEEDS_PLN`, never slug-cased.
Original framing, kept because it is still the rule for the NEXT stemless gig:
`POSTPROD.md` is explicit: *"A new gig is a copy of `judge_specs/<gig>.json` — never a
copy of a script."* That file is the **only** thing standing between CosmicFest and
the entire existing pipeline (`apply_boundaries``render_release`
`build_release_joins``build_release_plan` → the SC/YT/BC adapters). Both CosmicFest
deliverables sit downstream of this one file.
One known wrinkle to resolve while writing it: OPAL's spec assumes stems
(`stemsDir`, `stemmap`, `keeps`, `clipsDir`). CosmicFest has none. Either the spec
grows a stemless mode or the tools must tolerate absent stem fields — **verify, do
not assume**, per `feedback_verify_the_plumbing`.
### Open decisions that need PLN, not code
- **D-1 · the unreviewed merge `2f947c2`.** Still standing on master. PLN's own plan:
*"ill review migration diff and midiviz tomorrow"*. Nothing else should touch master
until this lands.
- **D-2 · A2 REVOLUTION mix fix vs the A1.7 signoff.** Cutting d5 requires a re-render,
and the board states plainly that any re-render **invalidates the ear signoff**.
Ship OPAL as signed, or re-cut REVOLUTION and re-get the ear on that one track?
- **D-3 · Is "Outro Dub Siren" (#14, 103 s) an album track?** PLN's note was a laugh —
*"pure dubsiren noise and fun ahahah it could be called 'Outro Dub Siren'"*. It
clearly belongs in the club mix. Whether it ships as a numbered track is a call.
- **D-4 · Soft edges — CORRECTED 2026-09-01.** The first draft of this line named
#10 and #12 from the *recovered tracklist's* unresolved list. Checked against
`setlist_ear.json`, which supersedes it, the real split is: **10 starts are
playhead calls** (#2-#11), #1 is the origin, and **#12 `LiveCode Parade`,
#13 `Vague de CRIME`, #14 `Outro: La Dub Sirène` were INFERRED** from notes on
skipped boundaries and have never been heard as cuts. #10 `PunkAChien` *was*
ear-called (*"perfect cut at 48:27.4 into PunkAChien!"*). Those **three** are the
audition list before any render is final.
- **D-5 · Rights gate the album split, not the club mix.** `rights_ledger.json`
re-measured: **103 banks — 43 third_party, 31 unknown, 19 dirt_samples, 10 cleared**.
PLN's rule stands: *"cant push the single tracks when samples hit, but full mix might
well pass!"* This is the argument for shipping **club mix first, split second** on
every public platform.
- **D-6 · Artwork.** Still none chosen for OPAL; the `output/*.jpg` set was rejected as
GenAI of unknown provenance. CosmicFest has none either.
### Collab thread — piment_bresilien
PLN: *"my friend wants to collab on one of the cosmicfest lineup tracks, i think
piment bres."* Confirmed present in **both** sets — OPAL track 07 (rendered, and
already uploaded to SC as `/piment-bresilien`) and CosmicFest #7 (2145.6→2525.4,
379.8 s, an ear-called playhead boundary: *"starting at 35:45.6 i hear proper only
piment start sound"*). Score lives at `live/collab/raph/piment_bresilien.tidal`
(103 lines, 12 orbits) and carries **one uncommitted edit** — line 33,
`-- |+ note 12.``|- note 12`, i.e. the melodic layer transposed down an octave
instead of up. Decide whether that edit is the collab's starting point before
branching.
### RESUME POINT — 2026-09-01 ~23:00, unattended run in flight
Three `systemd --user` units carry the CosmicFest pipeline. They **survive session
teardown** (verified: every harness background waiter in that session was killed while
all three units kept running). Check state before doing anything:
```bash
systemctl --user is-active cosmic-master.service cosmic-demucs.service cosmic-finish.service
journalctl --user -u cosmic-master.service -o cat | tail -20
cat /home/pln/Work/Sound/Prod/Cosmic26_master/FINISH_SUMMARY.txt # written by the finisher
```
`cosmic-finish.service` is ordered `After=` both producers and does split → verify →
stem EDA on its own, so the expected end state (~03:00) with **no further action** is:
- `Cosmic26_v1_streaming.flac` (+ `club` if its target was reachable)
- `tracks_v1_streaming/` — 14 FLACs, and `tracks_v1_club/` if the club master exists
- `eda_stems.json` + `master_report.json` + `FINISH_SUMMARY.txt`
- `stems_demucs/sections/01..14/{drums,bass,other,vocals}.wav` + `stems_demucs/global/`
**Progress metrics that LIE — do not reuse them.** Counting dirs under
`stems_demucs/sections/` counts sections *STARTED*, because a dir appears the moment the
writers open. And libsndfile buffers, so a stem sits at **44 bytes** (a bare WAV header)
through minutes of real work. Judge progress by `mtime` + flushed bytes, or just by
whether the unit is still `activating`.
#### The first question to ask of `eda_stems.json`
Q1 decides a mastering parameter that is currently a deliberate guess. The air shelf is
`AIR_G = 5.0` in `master_stemless.py` against a measured 8–16 kHz share of **0.12%**,
which would justify far more — held back because a 320 k MP3 puts codec hiss in that band
too. The EDA prints the verdict directly: if the HF is concentrated in `drums` by ≥2×, it
is cymbals and the shelf can open; if it is spread evenly across all four stems, it is
noise and the shelf stays shut. **A re-render on that basis invalidates nothing** — PLN
has not yet heard any of it.
#### Blocked on PLN, in priority order
1. **`content/lives/2026/cosmicfest-2026.md` + `cosmicfest-2026/tracks.json`** — the hard
gate in `build_release_plan.py`; without it NO upload adapter can run. Derivable parts
are drafted (`build_gig_tracksjson.py`); he owes **set title, venue, stage,
description** and the per-track `section` movements.
2. **The tone** on the streaming master — the one thing no measurement settles.
3. `Outro: La Dub Sirène` — album track or mix-only? (renders as `Outro_ La Dub
Sirène.flac`; the FLAC title tag keeps the colon.)
4. Club target: −9 LUFS may be unreachable here without flattening the music (the LRA
floor tripped on the trial at −10.5). Chase it or ship the miss?
5. **The unreviewed merge `2f947c2`** — still standing on master.
#### Ready to run, deliberately NOT done tonight
OPAL's SoundCloud has **7 of 15** tracks still missing and the board's re-run command is
idempotent by permalink. Left alone because PLN was actively sharing OPAL links that
evening and mutating the account under him was the wrong risk. Its two known defects
(`take-five-drops` uploaded ~24 s short; `ghosts-in-the-toilets` has one malformed tag)
still need the browser edit FORM — api-v2 PUT is dead for both create and update.
⚠️ **Do not upload CosmicFest's 14-track split before PLN approves the tone.** SoundCloud
is idempotent by permalink: a v2 re-render would be silently SKIPPED, not replaced. The
continuous mix goes up first, for judgement; the split follows approval.
# 038 — A record from one microphone
**2026-09-01.** PLN reframed the release: not one record but four deliverables —
{OPAL-26, CosmicFest-26} × {continuous club mix, per-track album split}. Then he
named the constraint himself: *"i ack cosmicset is a raw record from stage with
effects, we can only do our best, 80/20 likely"*, and later, flatly, *"we have
neither set action logs nor the stems.. we can interpolate from the code, but its
gonna be a educated guess game."*
This log is the night that turned that guess game into a pipeline.
## Where the two sets actually stood
Measured from the filesystem, not recalled from the board:
| | OPAL-26 | COSMICFEST-26 |
|---|---|---|
| source | 12 stems + master, 48 k/24 | **one stage-mic MP3**, 320 k, 62:59 |
| boundaries | ✅ ear-signed 08-16 | ✅ ear pass 08-29 |
| gig spec | ✅ | 🔴 **did not exist** |
| mix / master | ✅ v4 premix | 🔴 nothing |
| album split | ✅ 14 × 2 | 🔴 none |
| club mix | ✅ 2 × nodrop | 🔴 none |
| ear-reviewed MIX | ✅ A1.7 | ⛔ nothing existed to review |
So the honest answer to "are we done with boundaries, is the mix done, have we
reviewed it" was: **boundaries done for both; mix done and reviewed for OPAL
only; CosmicFest had never had a single mastering pass.** Its ear session settled
*where the cuts are*, never *how it sounds*.
## The thing that mattered most took two minutes
`ZOOM0067.MP3` — the only recording of a 63-minute gig — existed in exactly one
place: `~/Downloads`. Not on the freebox, not in `Prod/`. It is now in
`Prod/Cosmic26_master/` and mirrored to the freebox, `sha256 7a197589…d671b`
identical on both, and the Downloads original was deleted only after both copies
verified. A move, not a hopeful `mv`.
Adjacent trap, recorded because the names invite the error: `Prod/cosmicfest/`
and `Prod/cosmicfestv0.live1.*` are the **2025** edition, files dated
2025-06-27/28. Nothing in them belongs to this release.
## The tape is much better than "stage mic with FOH effects" implies
`ebur128` over the full 62:59: **I −25.2 LUFS · true peak −5.1 dBFS · LRA
11.1 LU**. Read as three pieces of good news: the peak says it was recorded
conservatively and is **not clipped**; 11.1 LU says the desk did **not crush** it
(a mangled feed lands under 6); it is simply *quiet*.
But the spectrum is the finding that reshaped the work:
| band | share | |
|---|---|---|
| 300–1000 Hz | **34.75%** | boxy midrange dominates |
| 1–4 kHz | 14.76% | |
| 4–8 kHz | **1.75%** (−17.6) | presence nearly absent |
| 8–16 kHz | **0.12%** (−29.4) | almost no air |
| <30 Hz | 0.07% | the HPF is free headroom |
L/R correlation 0.833 — genuinely stereo, so bass-mono below 120 Hz is safe.
That profile is a **microphone in a room**, not a desk feed: the PA's treble
never reaches the mic position, the room absorbs what does, mid-bass piles up.
**The tape is dark and boxy, not merely quiet** — so the dominant move is a
broadband *tilt*, and gain alone would have produced a louder muddy record. This
is the whole argument for measuring before reaching for a filter.
## The keystone was one file
`POSTPROD.md`: *"A new gig is a copy of `judge_specs/<gig>.json` — never a copy
of a script."* That file did not exist for CosmicFest, and **both** its
deliverables sat downstream of it. Written, and the stemless path **verified by
reading the code rather than by trying it and hoping**: `apply_boundaries` and
`render_release` touch no stem key at all — only `segments`, `master`,
`variants`, `releaseRoot`, `releaseTag`. The one tool that genuinely needs stems
is `build_judge_set`, which derives nominal segments from per-orbit activity.
`setlist_to_segments.py` fills exactly that hole, and its real job is keeping two
grades of truth apart. The setlist silently mixes them: some starts are playhead
calls PLN made while listening; others were reconstructed from his notes on
boundaries he *skipped*. Only the former enter `verified`, so `apply_boundaries`
prints `ear` or `nominal` per row and asserts every playhead call reached the
output. An inferred edge can never later be mistaken for one he heard.
**That distinction immediately corrected the board.** It had named #10 and #12 as
the soft edges, taken from the recovered tracklist's unresolved list. The ear file
supersedes it: ten starts (#2–#11) are playhead calls, #1 is the origin, and the
inferred ones are **#12 LiveCode Parade, #13 Vague de CRIME, #14 Outro: La Dub
Sirène**. #10 PunkAChien was ear-called all along — *"perfect cut at 48:27.4 into
PunkAChien!"*. Three tracks, not two, and a different three. The board also said
11 decided boundaries; there are 10.
Result: **14 tracks, 63.0 min of a 63.0 min master, zero gaps, all assertions
green.**
## Titles: what shipped, then PLN, never a slug
Resolution order is authored → OPAL-release → catalog → `NEEDS_PLN`. The
authoritative source for a shared track is what actually *shipped*: OPAL-26's
ear-signed `segments_v4.json` joined to that gig's `tracks.json` on performance
order, because neither file alone holds both the score path and the released
title. Four tracks were in no shipped release and no catalog entry, so the first
run emitted them as `NEEDS_PLN` rather than title-casing them into something
plausible. PLN then gave all four: **Rose Rouge · Mafia sans Serif · LiveCode
Parade · Outro: La Dub Sirène** (he asked for proper French accents). They live in
`judge_specs/cosmicfest-2026_titles.json` with value+source+locator+date.
That file also carries a `score` alias, because the gig slug `mafia` is
`mafia_sans_serif.tidal` — and the alias then **cross-checked itself**, yielding a
declared 160 BPM that matches OPAL's Mafia exactly. `Outro: La Dub Sirène` has no
`.tidal` at all, being a played outro rather than a composed track, which is also
why the catalog can never hold its title: the catalog is keyed by score path.
## Demucs: a lens for a gig with no stems
PLN: *"we can also use demucs on the raw recording, once we have boundaries, on
each section (does it make sense) then we have 'stems' to work with, its a good
signal imo"* — then *"demucs all 14 sections and a global take to see how both are
interpreted locally and globally"*.
It makes sense, and for a stronger reason than stated: `feedback_mastering_eda`
makes stem EDA a *required* first step, so without stems we could not follow our
own rule. Demucs restores the ability to measure. But one measured fact bounds
what the local/global comparison can show — **htdemucs' receptive field is 7.8 s**
(`model.segment = 39/5`) and it always splits internally, so a 63-minute file
gives the model no more musical context than a 4-minute one. The difference is
**input normalisation**, not context. On a set at −25.2 LUFS with 11.1 LU of
range that is still a real lever, but it is a gain-staging experiment and should
be described as one.
Three things cost real time and are worth keeping:
- **The demucs CLI cannot save here at all.** torchaudio 2.10 routes writes
through TorchCodec, which is not installed, so `python -m demucs` separates for
87 seconds and dies on `ImportError`. Going through `apply_model` + `soundfile`
fixes it and installs nothing into PLN's venv.
- **The first run was OOM-killed at 8.2 GB after 2 of 14 sections.** The
arithmetic was available beforehand and I did not do it: a 532 s section is
188 MB in, 750 MB out across four sources, and `apply_model`'s overlap
accumulator wants that again — ~2 GB × 4 workers, on a machine already holding
18 GB of unrelated Gradle builds. The fix was already written in the other half
of the same file: `global_pass` chunked its work for the same reason, so both
modes now share one bounded applier and differ only in which normalisation
stats they receive. Peak fell to 5.45 GB.
- **Two chunking details are load-bearing.** Seams are joined with an
equal-power ramp, because the overlap sums two *estimates of the same audio*
and a linear fade dips ~3 dB exactly at the join. And output is length-exact —
verified at 3969000 frames on all four stems for a 90 s input, zero drift —
which matters because the EDA seeks into the global stems by absolute time.
Validation before committing six hours: 30 s slice, four stems written, all
populated, sum-minus-original **27.4 dB below** the signal, and `vocals` at
−50.6 dBFS RMS on an instrumental passage — correctly near-empty. Agreement where
it should agree is the only thing that makes later disagreement worth trusting.
### Correction
Commit `37eb401` claims the chunked version was roughly twice as fast ("four
sections in the time two took"). **That claim is wrong and came from a broken
metric.** I counted directories under `stems_demucs/sections/`, but a directory
appears the moment `chunked_separate` opens its writers — so the count measures
sections *started*, not finished. libsndfile also buffers, so file size lags
badly: three of four in-flight sections sat at 44 bytes (a bare WAV header) after
nine minutes of real work. Measured properly via mtime and flushed bytes, one
worker was running at ~4.3× realtime under four-way contention plus the master
render. Chunking fixed the OOM; it did not demonstrably speed anything up.
## Mastering a bus with no stems
`master_stemless.py` is MASTERING.md's chain with the per-stem half removed and
one stage added because the measurement demanded it: HPF 25, a −2.5 dB dip at
450 Hz for the 34.75% pileup, +2 presence at 3.5 k, a +5 air shelf, LPF 19.5 k,
bass-mono below 120, the 1.5:1 glue comp, then two-pass `loudnorm` with
`linear=true`.
Restraint on the one number that begged for more: the air shelf is **+5, not the
+12 the deficit suggests**, because a 320 k MP3 puts codec residue in that band
alongside cymbals and a big boost lifts hiss into the master. Which of the two it
is, is precisely Q1 of the stem EDA — so the shelf stays modest until the lens
answers.
Bass-mono runs in **mid/side**, not by splitting and re-summing bands: a
`lowpass(120)` summed with a `highpass(120)` leaves a phase notch at the
crossover, whereas high-passing the side channel alone never filters the mid.
Two failures worth keeping, both about the loud target:
1. **Limiter before gain does nothing.** The signal reaching it still peaked at
−9.5 dBTP, so a −1 dBFS ceiling never engaged, and `linear=true` — correctly
refusing to breach TP — capped the master at −11.2 LUFS against −9. The gain
has to come first so the limiter has something to catch.
2. **One pass still undershoots**, because every dB the limiter absorbs is a dB
the final `loudnorm` cannot add. The shortfall *is* the missing gain, so
feeding it back converges — bounded at three attempts and guarded by an **LRA
floor**, since a loudness target reached by flattening the music is not
reached. On the 90 s trial: streaming **−14.0 LUFS, peak −1.0, 1.8 LU lost**;
club stopped at −10.5 with the floor tripped, reporting the miss rather than
clipping its way to a number.
## The gate that stops the shipment, and why it is right
`build_release_plan.py` refuses without canonical www gig metadata:
*"Gig metadata is never invented here — create it there first."* CosmicFest 2026
has no page under `content/lives/2026/`, and every upload adapter reads that
plan. So **no upload path is open** until PLN supplies the set title, venue and
stage. The gig spec leaves `album` absent for the same reason.
Rather than weaken the gate, `build_gig_tracksjson.py` draws its line explicitly:
*derivable* (name, file, bpm, timecodes, style from the score's directory,
samples from `catalog_view`'s parser) is generated; *PLN-only* (gig title, venue,
stage, description, and each track's `section`) is emitted as `null` with a
`_needs_pln` list. A draft, in the scratchpad, for him to promote. Tomorrow is
four fields, not an afternoon.
The derived timecodes then validated themselves against his own ear notes: Piment
Bresilien at **0:35:46** against *"starting at 35:45.6 i hear proper only piment
start sound"*, and LiveCode Parade at **0:54:58** against the 54:59 note. 14
tracks, 63.0 min, 30 sample packs, 89–170 BPM. Seven of fourteen scores are
absent from a stale `catalog_view`, so their sample lists are empty and *reported*
rather than silently blank.
**A second reason not to rush:** SoundCloud uploads are idempotent by permalink,
so publishing the 14-track split before PLN's ears approve the tone would poison
those permalinks — a v2 re-render is silently *skipped*, not replaced. The board
already records that defect for OPAL's `take-five-drops`.
## Two lessons about running work overnight
**Retracted, same night.** I first wrote here that `nohup … &` inside a tool call
does not survive, on the evidence of an empty log and no matching process. Both
were false negatives, and the wrong diagnosis cost more than a dead job would have.
The log was empty because **Python block-buffers stdout to a file** — no TTY, ~8 KB
buffer, and I had set `PYTHONUNBUFFERED=1` only in the systemd unit, not on the
`nohup` line. The process was missing because **rtk silently dropped the grep**;
`rtk proxy ps` showed it plainly, reparented to the user systemd manager and forty
minutes into its work. RTK.md says to fall back to `rtk proxy` when output looks
suspiciously empty, and I skipped that on the single check the whole conclusion
rested on.
So I started a second copy of the same 63-minute render. For ~40 minutes **two
processes wrote the same output paths**, racing on `Cosmic26_v1_streaming.flac` and
`master_report.json`. It surfaced only because two `ffmpeg` processes appeared at
different pipeline stages, which is a lucky tell rather than a designed one. Killed
the unsupervised duplicate by PID, deleted its 102 MB partial FLAC so the finisher
could not split a truncated file, and let the journaled unit continue.
Two rules earned the hard way: **an empty logfile is not evidence of no progress**
(use `python -u`), and **before starting a replacement, prove the original is
gone** — a duplicate writing the same paths is a corruption risk, not just waste.
The house rule survives intact and for its own reasons: systemd units outlive the
session. Verified the same night — every harness background waiter and monitor I
had armed was killed at a between-turn teardown, while all three `cosmic-*` units
kept running untouched.
And the steps *after* the long jobs were sitting in my session, which would have
left PLN with masters and nothing else. So `cosmic-finish.service` is ordered
`After=cosmic-master.service cosmic-demucs.service` and does the split, the
verify and the stem EDA unattended, writing `FINISH_SUMMARY.txt`. `Wants=` not
`Requires=`, so a club-target miss cannot cancel the split of a good streaming
master.
## Shipped
`39ca4de` board · `d97dda3` spec + converter + titles · `b8059f8` demucs +
mastering tools · `37eb401` the OOM fix · `6ae2b82` metadata drafter.
Branch `claude/release-board`.
## Still PLN's
- **The set title, venue and stage** — the only thing blocking every upload.
- **The tone**, on the streaming master: is +5 of air enough for a room this dark?
- Whether `Outro: La Dub Sirène` (103 s) is an album track or mix-only.
- Whether the club target is worth chasing below the LRA floor on this source.
- The unreviewed merge `2f947c2`, still standing on master.
---
## The results — 2026-09-02, ~02:00
### Delivered
| | |
|---|---|
| **Master A** `Cosmic26_v1_streaming.flac` | air shelf **+5 dB** · 661 MB |
| **Master B** `ab_air8/streaming.flac` | air shelf **+8 dB** · 673 MB |
| both | **−14.0 LUFS · −0.9 dBTP · LRA 7.5 / 7.6 · 44100 Hz / 24-bit** |
| duration | **3779.317551 s** against a source of 3779.317550 s |
| **album split** | `tracks_v1_streaming/` — 14 FLACs, verify **ALL OK**, format 44100/24 |
| demucs | 14 sections + a full 3779 s global pass |
Both of PLN's forms are therefore on disk: nothing was cut from this set, so the
continuous mix simply *is* the master file, and the split sits beside it.
The duration matching the source to within a microsecond is the load-bearing
number. It means every one of his playhead calls stays valid against the master,
so the split lands exactly where he put the cuts rather than approximately there.
### Q1 — the missing air is CYMBALS
Mean share of each stem's **own** energy above 8 kHz: drums **0.367%**, vocals
0.084%, other 0.033%, bass 0.000%. A **4.4× concentration in drums**, where evenly
spread energy would have meant codec hiss and a shelf that must stay shut.
So the deficit is real drum content the room ate, and opening the shelf is licensed
by measurement. How far is not, so the answer is an A/B rather than a number I
liked. The useful context for that choice: against the tape, 4–8 kHz went 1.75% →
3.14% (+5) → 3.86% (+8), and 8–16 kHz went 0.12% → 0.40% → 0.63%, against a rough
norm of 3–6% and 1–3% for electronic masters. **Both candidates are still below
typical for air.** The choice is "quite dark vs slightly less dark", not "safe vs
aggressive" — worth saying, because the framing changes the answer.
The A/B is also clean in the way that matters: identical loudness (−14.0), identical
true peak (−0.9), range within 0.1 LU. Only the tilt differs — 0.98× below 1 kHz,
1.23× at 4–8 k, 1.59× at 8–16 k. No "the louder one sounds better" confound.
### Q2 — the floor is worse here than OPAL
Kick active time, 40–120 Hz against each track's own p95: **mean 44.8%**, against
the **59%** `project_floor_problem` measured for OPAL-26 *from real orbit stems*.
Worst: **#2 There's Something About Drums 22.8%**, **#3 Am i Doing it Right 33.4%**,
**#8 Eh ouais je Funk 36.9%**. A stage mic in a room is a sufficient explanation on
its own, so this is a lead and not a verdict — but it is the first time this gig's
floor has been measurable at all.
### Q3 — the global pass was NOT worth running
PLN's own question, and the answer is a clean negative:
| stem | correlation | level delta |
|---|---|---|
| drums | 0.9765 | −0.14 dB |
| other | 0.9742 | −0.06 dB |
| vocals | 0.9543 | −0.17 dB |
| bass | **0.9432** | −0.21 dB |
**The two separations are the same stems.** The prediction held exactly: htdemucs'
7.8 s receptive field means a global pass cannot give the model more musical
context, so only the normalisation statistics differ — and they moved the level by
at most 0.21 dB. The residual 2–6% of uncorrelated energy is chunk-boundary
placement, not a different reading of the music.
It cost roughly 2.5 hours of CPU, competing with the master render, to establish a
≤0.21 dB difference. **Sections only on the next stemless gig.** The question was
worth asking once; it is not worth paying for twice. Banked in
`reference_demucs_on_xps22`.
Side finding: **bass is the least stably separated source** (0.9432, lowest of the
four), which is consistent with the same run putting PunkAChien's bass into
`other` — bass stem at −81.0 dBFS while `other` was the loudest stem in the set.
Anyone reading stem filenames as roles would have concluded that track has no bass.
### Three self-inflicted errors, same shape
All three were **checks whose failure mode was indistinguishable from success**:
1. **The nohup "death".** `rtk` silently dropped a `grep`, and Python block-buffered
a logfile to zero bytes. I concluded the job was dead and started a second copy;
two renders then wrote the same paths for ~40 minutes.
2. **The demucs speed claim.** Counted output directories, which appear when writers
*open*. Retracted from the commit, the log, and two memory files.
3. **A silent no-op edit.** A scripted replacement whose search string carried
escaped line-continuations that did not match the file, while the script printed
"gate tightened" unconditionally. Caught by grepping the file rather than
trusting the message.
And one that was not mine to invent but was mine to prevent: the first full master
came out at **192 kHz, 1.32 GB** because ffmpeg's `loudnorm` oversamples for
true-peak detection and never resamples back. `topic_postprod_mastering` has carried
that exact trap since tidal-ears hit it. I wrote a new mastering chain without
re-reading the omnibus that exists to prevent it. The fix is structural — the
resample lives in the chain *builder*, and the in-spec gate now asserts sample rate
and channel count beside I/TP/LRA, because a loudness-only gate certified a
4.3×-oversized master as ✓.
#!/usr/bin/env python3
"""build_gig_tracksjson — draft the canonical www `tracks.json` for a gig.
Why this exists: `build_release_plan.py` refuses to emit anything without
`Web/www/content/lives/<year>/<gig>.md` + `<gig>/tracks.json`, saying "Gig
metadata is never invented here — create it there first." Every upload adapter
reads that plan, so a missing gig page blocks SoundCloud, YouTube and Bandcamp
alike. CosmicFest 2026 has no page.
Most of that file is DERIVABLE and the rest is not, so this tool draws the line
explicitly instead of blurring it:
derivable name, file, bpm, start/end/duration, style, samples
PLN ONLY the gig's title, venue, stage, description (the .md frontmatter)
and each track's `section` (the set's movements)
Anything in the second group is emitted as `null` with a `_needs_pln` list, never
as a plausible guess — `feedback_metadata_vs_mastering` and
`feedback_metadata_provenance`. The output is a DRAFT written wherever you point
it; promoting it into the www repo is PLN's call, not this script's.
`samples` reuses `catalog_view.json`'s `score_sounds`, which is the existing
score parser. A track absent from that view is reported rather than given an
empty list, because an empty sample list is invisible once it is on the shelf
([[feedback_parsers_over_copy]]).
"""
from __future__ import annotations
import argparse
import json
import sys
from pathlib import Path
HERE = Path(__file__).resolve().parent
REPO = HERE.parent.parent # …/Sound/Tidal
sys.path.insert(0, str(REPO / "tools")) # sample_tfidf lives there
# The style OPAL published per track matches the score's own directory, which is
# how the corpus is organised. Derive it there rather than restating it.
# "collab" is in here because OPAL published it: every live/collab/* score came
# out as style "collab" in content/lives/2026/opal-festival-2026/tracks.json.
# style_of takes the LAST matching part, so live/collab/nova/techno/x still reads
# techno — collab only wins when nothing more specific is in the path.
STYLE_DIRS = {"techno", "dnb", "lounge", "acid", "jazz", "remix", "nujazz",
"breaks", "chip", "liquid", "house", "dub", "ambient", "collab"}
def hms(s: float) -> str:
s = int(round(s))
return f"{s // 3600}:{(s % 3600) // 60:02d}:{s % 60:02d}"
def style_of(score: str | None) -> str | None:
if not score:
return None
parts = [p for p in Path(score).parts if p in STYLE_DIRS]
return parts[-1] if parts else None
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("spec")
ap.add_argument("--out", required=True, help="draft tracks.json path")
ap.add_argument("--md-out", help="also draft the .md frontmatter skeleton")
ap.add_argument("--meta", help="JSON sidecar carrying the PLN-only fields, each "
"a bare value or {value, source, locator, date}")
a = ap.parse_args()
meta = json.loads(Path(a.meta).read_text()) if a.meta else {}
prov: dict = {}
def mval(key, default=None):
"""A supplied fact, with its provenance kept rather than dropped.
`feedback_metadata_provenance`: a value on its own cannot be audited six
months later, so a {value, source, locator, date} entry keeps the chain
into the output. A bare value is still accepted — sometimes PLN just says
it in chat — but it is recorded as unsourced so the gap stays visible.
"""
if key not in meta:
return default
v = meta[key]
if isinstance(v, dict) and "value" in v:
prov[key] = {k: v[k] for k in ("source", "locator", "date") if k in v}
return v["value"]
prov[key] = {"source": "unsourced — supplied bare in the meta sidecar"}
return v
# Movements, keyed by release track number. Same value-or-dict shape.
sections = {str(k): (v["value"] if isinstance(v, dict) and "value" in v else v)
for k, v in (meta.get("sections") or {}).items()}
if "sections" in meta:
prov["sections"] = (meta.get("_sections_provenance")
or {"source": "unsourced"})
spec = json.loads(Path(a.spec).read_text())
segs = json.loads(Path(spec["segmentsRelease"]).read_text())
nominal = {s["track"]: s for s in
json.loads(Path(spec["segments"]).read_text())}
view = json.loads((HERE / "catalog_view.json").read_text())["tracks"]
sounds = {t["track"]: t.get("score_sounds") or [] for t in view}
def parse_score(rel: str) -> list[str] | None:
"""Read the sounds straight out of the .tidal when the view has no row.
catalog_view's track universe comes from `load_gigs()`, which globs the
already-published `content/lives/**/tracks.json`. So a track making its
FIRST catalogued appearance — which is precisely what a new gig page is
full of — cannot be in the view yet: the file that would add it is the one
being written here. Rebuilding the view does not fix it; that is a cycle,
not staleness.
The fix is not a second parser. `feedback_one_parser_per_concept`: this
calls the same `tidal_score.orbit_sounds` the view itself calls, so a new
track's sample list is produced by the one authority on what a score
loads, and cannot drift from the catalogued ones.
"""
path = REPO / rel
if not path.exists():
return None
try:
import tidal_score as ts
from sample_tfidf import sound_vocab
vocab, kind = sound_vocab()
return sorted({d["sound"] for d in
ts.orbit_sounds(path, vocab, kind).values()})
except Exception as exc: # a broken parse is not a zero
print(f"⚠ score parse failed for {rel}: {exc}")
return None
rows, no_sounds, needs, first_appearance = [], [], [], []
for s in segs:
nom = nominal.get(s["perf_track"], {})
score = nom.get("score")
snd = sounds.get(score)
if score and snd is None:
snd = parse_score(score)
if snd is None:
no_sounds.append(score)
else:
first_appearance.append(score)
bpm = nom.get("bpm")
rows.append({
"name": s["title"],
"file": score,
"bpm": bpm if isinstance(bpm, (int, float)) else None,
"section": sections.get(str(s["track"])), # PLN only — the movements
"style": style_of(score),
"start_s": s["start"],
"end_s": s["end"],
"duration_s": s["duration"],
"start": hms(s["start"]),
"end": hms(s["end"]),
"samples": sorted(snd) if snd else [],
})
if rows[-1]["section"] is None:
needs.append(f"section for #{s['track']} {s['title']}")
still_needed = [k for k in ("title", "venue", "stage", "description")
if mval(k) in (None, "")] + needs
out = {
"_draft": [
"Generated by build_gig_tracksjson.py. Derived fields (name, file, bpm,",
"timecodes, style, samples) come from the gig spec, the ear-verified",
"segments and catalog_view.json's score parser. The fields that are NOT",
"derivable were supplied via --meta and their origin is in _provenance.",
"`_needs_pln` lists what is still genuinely missing — if it is empty,",
"nothing here is a guess.",
],
"gig": spec["gig"],
"title": mval("title"), # PLN only
"date": mval("date", spec["date"]),
"venue": mval("venue"), # PLN only
"stage": mval("stage"), # PLN only
"trackCount": len(rows),
"totalDuration_s": round(sum(r["duration_s"] for r in rows), 2),
"bpmRange": ([min(r["bpm"] for r in rows if r["bpm"]),
max(r["bpm"] for r in rows if r["bpm"])]
if any(r["bpm"] for r in rows) else None),
"samplePacks": sorted({x for r in rows for x in r["samples"]}),
"_needs_pln": still_needed,
"_provenance": prov,
"tracks": rows,
}
Path(a.out).write_text(json.dumps(out, indent=2, ensure_ascii=False) + "\n")
print(f"{'#':>3} {'start':>9} {'bpm':>5} {'style':<8} {'samples':>7} name")
for r, s in zip(rows, segs):
print(f"{s['track']:>3} {r['start']:>9} {str(r['bpm'] or '—'):>5} "
f"{str(r['style'] or '—'):<8} {len(r['samples']):>7} {r['name']}")
print(f"\n{len(rows)} tracks · {out['totalDuration_s']/60:.1f} min · "
f"{len(out['samplePacks'])} sample packs · bpm {out['bpmRange']}")
if first_appearance:
print(f"ℹ {len(first_appearance)} score(s) making their FIRST catalogued "
f"appearance — samples read straight from the .tidal: "
+ ", ".join(Path(p).stem for p in first_appearance))
if no_sounds:
print(f"⚠ {len(no_sounds)} score(s) with NO sample list — not in "
f"catalog_view and not parseable on disk: "
+ ", ".join(Path(p).stem for p in no_sounds))
if still_needed:
print(f"⚠ NEEDS PLN ({len(still_needed)}): " + " · ".join(still_needed[:6])
+ (" …" if len(still_needed) > 6 else ""))
else:
print("✓ nothing left to invent — every non-derivable field is sourced")
if a.md_out:
Path(a.md_out).write_text(
"---\n"
+ ("" if not still_needed else
f'# DRAFT — {len(still_needed)} field(s) still PLN\'s to fill.\n')
+ f'title: "{out["title"] or "TODO set title"}"\n'
f'date: "{out["date"]}"\n'
f'time: "{mval("time", "")}"\n'
f'location: "{mval("location", "TODO")}"\n'
f'address: "{out["venue"] or "TODO venue"}"\n'
f'stage: "{out["stage"] or "TODO"}"\n'
f'description: "{mval("description", "TODO")}"\n'
f'ctaURL: "{mval("ctaURL", "")}"\n'
f'ctaText: "{mval("ctaText", "")}"\n'
f'video: ""\naudio: ""\narchive: ""\n'
f'tags: {json.dumps(mval("tags", ["livecoding", "tidalcycles"]), ensure_ascii=False)}\n'
"---\n\n"
f"# {out['title'] or 'TODO set title'}\n\n"
f"{len(rows)} morceaux, {out['totalDuration_s']/60:.0f} minutes. "
"La tracklist complète, avec les timecodes mesurés sur l'enregistrement, "
"est dans `tracks.json`.\n")
print(f"✓ {a.md_out} (frontmatter skeleton)")
print(f"✓ {a.out}")
return 0
if __name__ == "__main__":
sys.exit(main())
......@@ -33,6 +33,7 @@ from __future__ import annotations
import argparse
import json
import re
import sys
from datetime import datetime, timezone
from pathlib import Path
......@@ -94,15 +95,22 @@ def check_signoff(ear: dict, tracks: list[Path], force: bool) -> list[str]:
problems.append(
"no `release_signoff` in the ear file — nobody has confirmed this "
"record by ear. Run render_release + build_release_joins and listen.")
return problems
signed = datetime.fromisoformat(so["date"]).date()
newest = max((datetime.fromtimestamp(p.stat().st_mtime).date() for p in tracks),
default=None)
if newest and newest > signed:
problems.append(
f"the rendered tracks ({newest}) are NEWER than the signoff "
f"({signed}). Audio has changed since PLN approved it; re-audition "
f"before uploading.")
# NO early return. This branch used to `return problems` right here,
# which meant --force — documented as "emit even without a valid ear
# signoff" — could only ever override the STALE-signoff problem, never
# the missing-signoff one it was named for. The first gig to need it
# (CosmicFest 2026: private audition upload, because SoundCloud is where
# PLN listens, so the signoff cannot precede the upload) is how that came
# out. Fall through to the shared override below.
else:
signed = datetime.fromisoformat(so["date"]).date()
newest = max((datetime.fromtimestamp(p.stat().st_mtime).date() for p in tracks),
default=None)
if newest and newest > signed:
problems.append(
f"the rendered tracks ({newest}) are NEWER than the signoff "
f"({signed}). Audio has changed since PLN approved it; re-audition "
f"before uploading.")
if force and problems:
for p in problems:
print(f" ! OVERRIDDEN: {p}", file=sys.stderr)
......@@ -116,6 +124,13 @@ def main() -> int:
ap.add_argument("spec")
ap.add_argument("--variant", default="streaming")
ap.add_argument("--artwork", help="cover image applied to every track")
ap.add_argument("--permalink-suffix", default="",
help="append to every permalink, e.g. `cosmicfest-2026`. A "
"live set's tracks are NOT unique on the shelf: the same "
"score gets played at several gigs, and SoundCloud is "
"idempotent by permalink, so the second gig's upload "
"either silently skips or overwrites the first's. Scope "
"the permalinks per gig and neither can happen.")
ap.add_argument("--include-mix", action="store_true",
help="also publish the continuous *_nodrop mix, listed first")
ap.add_argument("--out", required=True)
......@@ -123,7 +138,11 @@ def main() -> int:
help="track names from the canonical tracklist (default) or "
"from the rendered filenames")
ap.add_argument("--force", action="store_true",
help="emit even without a valid ear signoff (say why in the log)")
help="emit even without a valid ear signoff (say why in --why)")
ap.add_argument("--why", default="",
help="required with --force: why this plan may exist unsigned. "
"Recorded IN the plan, because a terminal warning scrolls "
"away and the JSON is what the uploader reads.")
a = ap.parse_args()
spec = json.loads(Path(a.spec).read_text())
......@@ -146,6 +165,10 @@ def main() -> int:
if len(files) != len(segs):
sys.exit(f"{len(files)} rendered files but {len(segs)} segments in {tdir}")
if a.force and not a.why.strip():
sys.exit("--force needs --why. An override with no stated reason is "
"indistinguishable from an approved release once the terminal "
"is closed, and this plan is what the uploader trusts.")
problems = check_signoff(ear, files, a.force)
if problems:
print("REFUSING TO WRITE A PLAN — the record is not ear-approved:")
......@@ -170,6 +193,10 @@ def main() -> int:
print(f" {t['name']!r}")
return 1
def permalink(title: str) -> str:
base = re.sub(r"-+", "-", re.sub(r"[^a-z0-9]+", "-", title.lower())).strip("-")
return f"{base}-{a.permalink_suffix}" if a.permalink_suffix else base
tracks, renamed = [], []
for seg, f in zip(segs, files):
src = by_name[match_key(seg["title"])]
......@@ -187,6 +214,7 @@ def main() -> int:
tags = base_tags + [t for t in (src.get("style"), src.get("section")) if t]
tracks.append({
"title": title,
"permalink": permalink(title),
"audio": str(f),
"artwork": a.artwork,
"genre": src.get("style", ""),
......@@ -235,6 +263,9 @@ def main() -> int:
# Composed from canonical FIELDS (title · venue · date), never from
# a remembered string. The format is here; the facts are theirs.
"title": f"{album} — {fm.get('address','')} {fm.get('date','')[:4]} (full set)",
# No year in the base: the suffix already carries the gig, and
# "cosmicset-2026-2026-full-set-cosmicfest-2026" is nobody's URL.
"permalink": permalink(f"{album} full set"),
"audio": str(mix),
"artwork": a.artwork,
"genre": "livecoding",
......@@ -259,6 +290,16 @@ def main() -> int:
],
"album": album,
"artist": spec.get("artist", "ParVagues"),
# An unsigned plan must announce itself. The uploader's --sharing flag
# decides private vs public, so nothing here can enforce it; what this
# CAN do is make sure a plan built for a private audition is never
# mistaken for an approved record when someone re-runs it later.
"_UNSIGNED": ({"reason": a.why,
"forced": datetime.now(timezone.utc).isoformat(timespec="seconds"),
"means": "NOT ear-approved. Private audition only — do not "
"flip these permalinks public without a "
"release_signoff in the ear file."}
if a.force else None),
"date": fm.get("date"),
"_provenance": {
"album/date/description/tags": str(md_path.relative_to(WWW)),
......
This source diff could not be displayed because it is too large. You can view the blob instead.
#!/usr/bin/env python3
"""check_permalinks — refuse a release plan whose permalinks are already taken.
## Why this is a gate and not a warning
SoundCloud is idempotent by permalink, and the uploader treats an existing
permalink as "already done". That is the right behaviour for resuming a
half-finished upload and the wrong behaviour for a *different* recording of the
same track — and a live catalogue is full of those. ParVagues played Piment
Brésilien at Opal in August and at CosmicFest a fortnight later; both render to
`/piment-bresilien`.
So the failure mode is not an error. It is a release that reports success with
four tracks quietly missing, or — with `--overwrite` — the previous gig's
release silently replaced by this one's. Both are invisible from the terminal
that ran the upload. Caught on CosmicFest-2026 before `--go`: four of fourteen
permalinks were already Opal's private uploads.
## What it does
Resolves the plan's permalinks (explicit `permalink`, else the same slugify the
uploader applies to `title`) against the live account, and separates:
* FREE — nothing there, a normal upload
* OURS-RESUME — taken, and the duration matches this plan's audio, so it is
almost certainly a resumed upload of THIS release
* COLLISION — taken by audio of a different length: a different recording
A COLLISION is fatal unless `--allow-collisions`. Duration is the discriminator
because it is the one property that survives SoundCloud's transcode and is
already in the plan (`_duration_s`).
python3 check_permalinks.py <plan.json> [--tolerance 3]
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
sys.path.insert(0, "/home/pln/Work/Sound/tidal-ears/src")
def slugify(title: str) -> str:
return re.sub(r"-+", "-", re.sub(r"[^a-z0-9]+", "-", title.lower())).strip("-")
def shelf() -> dict[str, dict]:
from tidal_ears.sc import SC, resolve_creds
api = SC(resolve_creds())
me = api.get("/me")
out, path = {}, f"/users/{me['id']}/tracks"
# `access=playable,preview,blocked` is load-bearing: scform.py's own notes
# record that without it the owner's PRIVATE tracks are omitted, which would
# make every private permalink look free — the exact opposite of this gate's
# job, since an audition release is private by definition.
r = api.get(path, limit=200, access="playable,preview,blocked")
for t in (r.get("collection", []) if isinstance(r, dict) else r):
out[t["permalink"]] = {"title": t.get("title"),
"dur_s": round((t.get("duration") or 0) / 1000, 1),
"sharing": t.get("sharing")}
return out
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("plan")
ap.add_argument("--tolerance", type=float, default=3.0,
help="seconds of duration difference still counted as the "
"same audio (transcode + the uploader's own trim)")
ap.add_argument("--allow-collisions", action="store_true",
help="proceed anyway — say why in the release log")
a = ap.parse_args()
plan = json.loads(Path(a.plan).read_text())
live = shelf()
free, resume, collide = [], [], []
for t in plan["tracks"]:
pl = t.get("permalink") or slugify(t["title"])
got = live.get(pl)
if not got:
free.append((pl, t["title"]))
elif abs(got["dur_s"] - (t.get("_duration_s") or -1)) <= a.tolerance:
resume.append((pl, t["title"], got))
else:
collide.append((pl, t["title"], got, t.get("_duration_s")))
def hms(s):
s = int(s or 0)
return f"{s // 60}:{s % 60:02d}"
print(f"{plan.get('album')} · {len(plan['tracks'])} planned upload(s)")
print(f" {len(free):>3} free · {len(resume):>3} already up as this same audio "
f"· {len(collide):>3} COLLISION(S)")
if resume:
print("\nalready up, same length — an upload run would resume, not duplicate:")
for pl, title, got in resume:
print(f" /{pl:<44} {hms(got['dur_s'])} {got['sharing']}")
if collide:
print("\nCOLLISIONS — the permalink is taken by DIFFERENT audio:")
for pl, title, got, want in collide:
print(f" /{pl:<44} shelf {hms(got['dur_s'])} ({got['title']!r}) "
f"vs plan {hms(want)} ({title!r})")
print("\n An upload would SKIP these tracks (release silently short) or,")
print(" with --overwrite, replace the other release's audio. Re-generate")
print(" the plan with `build_release_plan.py --permalink-suffix <gig>`.")
if not a.allow_collisions:
return 1
print("\n ! OVERRIDDEN by --allow-collisions")
print("\n✓ no collisions" if not collide else "")
return 0
if __name__ == "__main__":
sys.exit(main())
#!/usr/bin/env python3
"""compare_masters — put two or more finished masters side by side, in numbers.
Built for an A/B that PLN has to judge by ear. The ear is the authority and this
does not try to replace it — but "which one do you prefer" is a much easier
question when you also know WHAT differs and by how much. A tilt of a few dB in a
shelf changes more than the band it names: it moves true peak, it changes how hard
the limiter works, and it shifts where the loudness normaliser lands.
So this reports, per file: integrated loudness, true peak, loudness range, and the
band-energy distribution — then a delta table against the first file, because the
differences are the point and eyeballing two absolute tables is how you miss a
0.3% change that matters.
Deliberately measured on the RENDERED files rather than predicted from the filter
settings. `feedback_verify_own_renders`: the machine's job is to catch objective
facts so PLN's ears are spent on taste. A shelf set to +3 dB does not necessarily
put +3 dB in the master — the limiter downstream has opinions.
"""
from __future__ import annotations
import argparse
import re
import subprocess
import sys
from pathlib import Path
import numpy as np
import soundfile as sf
BANDS = [(20, 60), (60, 120), (120, 300), (300, 1000),
(1000, 4000), (4000, 8000), (8000, 16000), (16000, 22050)]
def ebur128(path: Path) -> dict:
p = subprocess.run(["ffmpeg", "-nostdin", "-hide_banner", "-i", str(path),
"-af", "ebur128=peak=true", "-f", "null", "-"],
capture_output=True, text=True)
tail = p.stderr[-2000:]
def grab(label):
m = re.search(rf"{label}:\s*(-?[\d.]+)", tail)
return float(m.group(1)) if m else None
return {"I": grab("I"), "LRA": grab("LRA"), "peak": grab("Peak")}
def bands(path: Path) -> tuple[dict, float]:
"""Band shares over the whole file, streamed — these are hour-long masters."""
n = 1 << 15
win = np.hanning(n)
acc = np.zeros(len(BANDS))
corr_num = corr_l = corr_r = 0.0
with sf.SoundFile(str(path)) as f:
sr = f.samplerate
fr = np.fft.rfftfreq(n, 1 / sr)
idx = [(fr >= lo) & (fr < hi) for lo, hi in BANDS]
while True:
blk = f.read(n, dtype="float32", always_2d=True)
if blk.shape[0] < n:
break
L, R = blk[:, 0], blk[:, 1]
corr_num += float((L * R).sum())
corr_l += float((L * L).sum())
corr_r += float((R * R).sum())
S = np.abs(np.fft.rfft(blk.mean(axis=1) * win)) ** 2
for j, m in enumerate(idx):
acc[j] += float(S[m].sum())
tot = acc.sum() or 1.0
denom = (corr_l * corr_r) ** 0.5
return ({f"{lo}-{hi}": 100 * v / tot for (lo, hi), v in zip(BANDS, acc)},
corr_num / denom if denom > 0 else 0.0)
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("files", nargs="+")
ap.add_argument("--label", action="append", default=[],
help="label per file, in order (default: the filename)")
a = ap.parse_args()
paths = [Path(f) for f in a.files]
for p in paths:
if not p.exists():
print(f"MISSING {p}", file=sys.stderr)
return 1
labels = [a.label[i] if i < len(a.label) else paths[i].stem
for i in range(len(paths))]
rows = []
for p in paths:
lo = ebur128(p)
bs, corr = bands(p)
rows.append({"loud": lo, "bands": bs, "corr": corr,
"size_mb": p.stat().st_size / 1e6})
w = 13
print(f"{'':<14}" + "".join(f"{l[:w]:>{w}}" for l in labels))
print("-" * (14 + w * len(labels)))
for key, fmt in (("I", "{:.1f} LUFS"), ("peak", "{:.1f} dBTP"), ("LRA", "{:.1f} LU")):
print(f"{key:<14}" + "".join(
f"{(fmt.format(r['loud'][key]) if r['loud'][key] is not None else '—'):>{w}}"
for r in rows))
print(f"{'L/R corr':<14}" + "".join(f"{r['corr']:>{w}.3f}" for r in rows))
print(f"{'size':<14}" + "".join(f"{r['size_mb']:>{w-3}.0f} MB" for r in rows))
print(f"\nband energy share (%)")
print(f"{'':<14}" + "".join(f"{l[:w]:>{w}}" for l in labels))
for band in rows[0]["bands"]:
print(f"{band:<14}" + "".join(f"{r['bands'][band]:>{w}.3f}" for r in rows))
if len(rows) > 1:
print(f"\nDELTA vs {labels[0]} — the differences are the point")
print(f"{'':<14}" + "".join(f"{l[:w]:>{w}}" for l in labels[1:]))
base = rows[0]
for key in ("I", "peak", "LRA"):
if base["loud"][key] is None:
continue
print(f"{key:<14}" + "".join(
f"{r['loud'][key] - base['loud'][key]:>+{w}.2f}" for r in rows[1:]))
for band in base["bands"]:
b0 = base["bands"][band]
cells = []
for r in rows[1:]:
d = r["bands"][band]
# ratio, not difference: a band at 0.12% going to 0.24% is a
# doubling, and "+0.12 points" hides that completely.
cells.append(f"{d / b0:>{w}.2f}x" if b0 > 1e-9 else f"{'—':>{w}}")
print(f"{band:<14}" + "".join(cells))
return 0
if __name__ == "__main__":
sys.exit(main())
#!/usr/bin/env python3
"""demucs_sections — separate a stemless gig into per-track stems, two ways.
A set captured on one stage mic has no stems, so `feedback_mastering_eda`'s rule
("EDA on stems is a required first step; never infer a sound's role from its
name") is unsatisfiable. Demucs restores the ability to MEASURE. Read the output
as a lens, not as mix elements: re-summing separated stems gives the original
plus separation error, which stays masked for a small corrective move and turns
phasey for a real rebalance.
PLN asked for both a per-section pass and a global one, "to see how both are
interpreted locally and globally". One measured fact shapes what that comparison
can possibly show: **htdemucs' receptive field is 7.8 s** (`model.segment = 39/5`)
and it always splits internally, so the model NEVER sees musical context beyond
those 7.8 s no matter how long the file is. The local/global difference is
therefore NOT context — it is INPUT NORMALISATION. Demucs normalises by the
mean/std of whatever you hand it, so a section normalises against itself while a
global pass normalises against the whole set. On this gig that is a big lever,
not a nitpick: the set sits at -25.2 LUFS with 11.1 LU of range, so the quiet
sections get far less gain under global stats than under their own.
So `global` mode computes the normalisation over the WHOLE recording and then
applies the model in bounded chunks with those stats. That isolates the one
variable that can actually differ, and it does so in ~1 GB instead of the ~12 GB
a single 63-minute tensor would need (4 sources x 1.33 GB output, plus
apply_model's overlap accumulator, against 12 GB free). Chunks overlap and are
crossfaded so the seams the chunking introduces do not land in the audio.
Resumable by construction: an existing, non-empty stem set is skipped, because a
6-hour CPU job that cannot be restarted is a job that will be restarted from zero.
"""
from __future__ import annotations
import argparse
import json
import os
import subprocess
import sys
import time
from concurrent.futures import ProcessPoolExecutor
from pathlib import Path
SR = 44100
SOURCES = ("drums", "bass", "other", "vocals")
def slug(title: str) -> str:
keep = "".join(c if c.isalnum() or c in " -_" else "" for c in title)
return "_".join(keep.split()).lower()[:48]
def cut(master: Path, start: float, end: float, dest: Path) -> Path:
"""Decode one section to WAV. Demucs on the MP3 directly would re-decode the
whole file per section, and a WAV is what the analysis pass wants anyway."""
if dest.exists() and dest.stat().st_size > 1000:
return dest
dest.parent.mkdir(parents=True, exist_ok=True)
subprocess.run(["ffmpeg", "-nostdin", "-v", "error", "-y",
"-ss", f"{start:.3f}", "-to", f"{end:.3f}", "-i", str(master),
"-ar", str(SR), "-ac", "2", "-c:a", "pcm_s24le", str(dest)],
check=True)
return dest
def _load_model(threads: int):
import torch
from demucs.pretrained import get_model
torch.set_num_threads(threads)
m = get_model("htdemucs")
m.eval()
return m
def stats(path: Path) -> tuple[float, float]:
"""Streamed mean/std of the mono sum. Streamed because the whole point is to
never hold a long file in RAM (see chunked_separate's docstring)."""
import soundfile as sf
n, s1, s2 = 0, 0.0, 0.0
with sf.SoundFile(str(path)) as f:
while (blk := f.read(1 << 20, dtype="float32", always_2d=True)).size:
mono = blk.mean(axis=1).astype("float64")
n += mono.size
s1 += float(mono.sum())
s2 += float((mono ** 2).sum())
if not n:
return 0.0, 1.0
mu = s1 / n
return mu, max((s2 / n - mu * mu) ** 0.5, 1e-9)
def chunked_separate(src: Path, out_dir: Path, mu: float, sd: float,
chunk_s: float, xfade_s: float, model) -> None:
"""Separate `src` in bounded chunks, writing the four stems incrementally.
ONE code path serves both modes, because the modes differ ONLY in which
normalisation statistics they are handed — that is the whole local/global
distinction (htdemucs' receptive field is 7.8 s, so neither mode gives the
model more musical context). Sections pass their own stats, the global pass
passes the whole recording's.
Chunking is not an optimisation, it is the fix for a crash. The first
version handed each worker its entire section and four workers peaked at
8.2 GB, which the kernel OOM-killer ended after two of fourteen sections:
a 532-second section is 188 MB in, 750 MB of output across four sources,
and apply_model's overlap accumulator wants that again. Chunked, a worker
holds a few hundred MB no matter how long the track is.
Chunks overlap by `xfade_s` and are joined with an equal-power ramp: the
seam sums two estimates of the same audio, so a linear fade would dip 3 dB
exactly where the crossfade sits. Verified length-exact — a 90 s input came
back as 3969000 frames on all four stems, zero drift, which matters because
the EDA seeks into the global stems by absolute time.
"""
import math
import soundfile as sf
import torch
out_dir.mkdir(parents=True, exist_ok=True)
mu_t, sd_t = torch.tensor(mu, dtype=torch.float32), torch.tensor(sd, dtype=torch.float32)
with sf.SoundFile(str(src)) as f:
sr, total = f.samplerate, f.frames
chunk, xf = int(chunk_s * sr), int(xfade_s * sr)
writers = {n: sf.SoundFile(str(out_dir / f"{n}.wav"), "w", samplerate=sr,
channels=2, subtype="PCM_24") for n in model.sources}
try:
tail, pos = None, 0
while pos < total:
f.seek(pos)
want = min(chunk + xf, total - pos)
data = f.read(want, dtype="float32", always_2d=True)
if data.shape[0] < 1:
break
wav = torch.from_numpy(data.T).contiguous()
out = separate(wav, model, ref_mean=mu_t, ref_std=sd_t)
keep = min(chunk, out.shape[-1])
body, nxt = out[:, :, :keep].clone(), out[:, :, keep:].clone()
del out, wav
if tail is not None and tail.shape[-1]:
k = min(tail.shape[-1], body.shape[-1])
if k:
r = torch.linspace(0, 1, k)
body[:, :, :k] = (tail[:, :, :k] * torch.cos(r * math.pi / 2)
+ body[:, :, :k] * torch.sin(r * math.pi / 2))
for name, tensor in zip(model.sources, body):
writers[name].write(tensor.numpy().T)
tail = nxt if nxt.shape[-1] else None
pos += keep
finally:
for w in writers.values():
w.close()
def separate(wav, model, ref_mean=None, ref_std=None):
"""Run the model. `ref_*` override the per-input normalisation stats — that
override IS the whole local/global distinction."""
import torch
from demucs.apply import apply_model
ref = wav.mean(0)
mu = ref.mean() if ref_mean is None else ref_mean
sd = ref.std() if ref_std is None else ref_std
with torch.no_grad():
out = apply_model(model, ((wav - mu) / sd)[None],
device="cpu", shifts=0, split=True,
overlap=0.25, progress=False, num_workers=0)[0]
return out * sd + mu
def do_section(args) -> tuple[str, float, bool]:
"""One section, in its own process — torch threads do not scale to 16 on a
single job, so parallel sections beat one wide job."""
wav_path, out_dir, threads, chunk_s, xfade_s = args
wav_path, out_dir = Path(wav_path), Path(out_dir)
if all((out_dir / f"{s}.wav").exists() and (out_dir / f"{s}.wav").stat().st_size > 1000
for s in SOURCES):
return (out_dir.name, 0.0, True)
t0 = time.time()
mu, sd = stats(wav_path) # LOCAL stats: the section normalises to itself
model = _load_model(threads)
chunked_separate(wav_path, out_dir, mu, sd, chunk_s, xfade_s, model)
return (out_dir.name, time.time() - t0, False)
def global_pass(master: Path, out_dir: Path, chunk_s: float, xfade_s: float, threads: int):
"""Global normalisation, same chunked applier."""
out_dir.mkdir(parents=True, exist_ok=True)
if all((out_dir / f"{s}.wav").exists() and (out_dir / f"{s}.wav").stat().st_size > 1000
for s in SOURCES):
print("global: already done, skipping", flush=True)
return
print("global: pass 1/2 — normalisation stats over the FULL recording", flush=True)
mu, sd = stats(master)
print(f"global: mean={mu:.6g} std={sd:.6g}", flush=True)
model = _load_model(threads)
print("global: pass 2/2 — chunked separation with global stats", flush=True)
chunked_separate(master, out_dir, mu, sd, chunk_s, xfade_s, model)
print("global: done", flush=True)
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("spec")
ap.add_argument("--mode", choices=("sections", "global", "both"), default="both")
ap.add_argument("--workers", type=int, default=4, help="parallel sections")
ap.add_argument("--threads", type=int, default=4, help="torch threads per section")
ap.add_argument("--chunk-s", type=float, default=120.0,
help="bounded work unit; 4 workers x 120 s stays well "
"under the 8.2 GB peak that got OOM-killed")
ap.add_argument("--xfade-s", type=float, default=2.0)
a = ap.parse_args()
spec = json.loads(Path(a.spec).read_text())
master = Path(spec["master"])
root = Path(spec["releaseRoot"])
segs = json.loads(Path(spec["segmentsRelease"]).read_text())
sect_wav = root / "sections_wav"
stems = root / "stems_demucs"
if a.mode in ("sections", "both"):
jobs = []
for s in segs:
name = f"{s['track']:02d}-{slug(s['title'])}"
w = cut(master, s["start"], s["end"], sect_wav / f"{name}.wav")
jobs.append((str(w), str(stems / "sections" / name), a.threads,
a.chunk_s, a.xfade_s))
print(f"sections: {len(jobs)} jobs, {a.workers} workers x {a.threads} threads",
flush=True)
t0 = time.time()
with ProcessPoolExecutor(max_workers=a.workers) as ex:
for name, dt, skipped in ex.map(do_section, jobs):
print(f" {'skip' if skipped else 'done'} {name}"
+ ("" if skipped else f" {dt:.0f}s"), flush=True)
print(f"sections: {time.time()-t0:.0f}s total", flush=True)
if a.mode in ("global", "both"):
global_pass(master, stems / "global", a.chunk_s, a.xfade_s, min(16, os.cpu_count() or 8))
return 0
if __name__ == "__main__":
sys.exit(main())
#!/usr/bin/env python3
"""eda_stems — the mastering EDA a stemless gig could not have, from demucs stems.
`feedback_mastering_eda` makes EDA on stems a required first step and forbids
inferring a sound's role from its name. A single stage-mic file makes both
impossible. Demucs restores the ability to measure, so this is that measurement —
and it is a LENS, not a mix: `feedback_label_and_lens` still applies, so demucs'
source names ("vocals") are treated as hypotheses, and presence is gated before
any stem is compared to another (`presence_is_a_precondition`).
Three questions it exists to answer:
1. **Is the 8-16 kHz deficit cymbals or codec noise?** The master measured 0.12%
of total energy up there and the air shelf was held to +5 dB because a 320 k
MP3 puts hiss in that band too. If the HF lives in `drums`, it is cymbals and
the shelf can open up; if it is spread evenly across all four stems, it is
noise and the shelf should stay shut. A broadband boost cannot tell them
apart — only a per-stem view can.
2. **Where is the kick?** `project_floor_problem` measured the kick audible in
only 59% of OPAL-26. That was computed from real orbit stems. This is the same
question asked of a gig that has none, via 40-120 Hz activity over time in the
drums stem — reported as a percentage of each track, so a track whose floor
drops out is named rather than averaged away.
3. **Local vs global: how much does normalisation actually change?** htdemucs'
receptive field is 7.8 s (`model.segment = 39/5`) and it always splits
internally, so a global pass cannot give the model more musical context. The
only thing that differs is the normalisation statistics. This measures whether
that difference is audible-scale or a rounding error, per stem and per
section, by comparing each section's stem against the same span cut out of the
global stem.
"""
from __future__ import annotations
import argparse
import json
import sys
from pathlib import Path
import numpy as np
import soundfile as sf
SOURCES = ("drums", "bass", "other", "vocals")
BANDS = [(20, 60), (60, 120), (120, 300), (300, 1000),
(1000, 4000), (4000, 8000), (8000, 16000)]
PRESENCE_DBFS = -60.0 # below this a stem is empty, not quiet
KICK_LO, KICK_HI = 40, 120
KICK_WIN_S = 0.25
def db(x: float) -> float:
return 20 * np.log10(max(float(x), 1e-12))
def read(path: Path):
d, sr = sf.read(str(path), dtype="float32", always_2d=True)
return d.mean(axis=1), sr # mono for analysis; roles are not stereo
def band_shares(mono: np.ndarray, sr: int) -> dict:
n = 1 << 15
acc = np.zeros(len(BANDS))
win = np.hanning(n)
for i in range(0, max(len(mono) - n, 1), n):
seg = mono[i:i + n]
if len(seg) < n:
break
S = np.abs(np.fft.rfft(seg * win)) ** 2
fr = np.fft.rfftfreq(n, 1 / sr)
for j, (lo, hi) in enumerate(BANDS):
acc[j] += float(S[(fr >= lo) & (fr < hi)].sum())
tot = acc.sum() or 1.0
return {f"{lo}-{hi}": round(100 * v / tot, 2) for (lo, hi), v in zip(BANDS, acc)}
def kick_activity(mono: np.ndarray, sr: int) -> dict:
"""Fraction of time the 40-120 Hz band is within 12 dB of its own p95.
Relative to the track, not to an absolute dBFS: these stems come from a
quiet, room-coloured source, so any fixed threshold measures the recording
level rather than the kick. `feedback_measure_the_time_axis` — an aggregate
cannot tell a fade from a sparse pattern, so this walks the axis.
"""
from numpy.fft import rfft, irfft
n = len(mono)
F = rfft(mono)
fr = np.fft.rfftfreq(n, 1 / sr)
F[(fr < KICK_LO) | (fr > KICK_HI)] = 0
band = irfft(F, n=n)
w = max(int(KICK_WIN_S * sr), 1)
trim = (len(band) // w) * w
if trim == 0:
return {"active_pct": 0.0, "windows": 0}
env = np.sqrt((band[:trim].reshape(-1, w) ** 2).mean(axis=1))
if not env.size or env.max() <= 0:
return {"active_pct": 0.0, "windows": 0}
ref = np.percentile(env, 95)
thr = ref * (10 ** (-12 / 20))
return {"active_pct": round(100 * float((env >= thr).mean()), 1),
"windows": int(env.size),
"p95_dbfs": round(db(ref), 1)}
def stem_row(path: Path) -> dict | None:
if not path.exists() or path.stat().st_size < 1000:
return None
mono, sr = read(path)
rms, pk = float(np.sqrt((mono ** 2).mean())), float(np.abs(mono).max())
row = {"rms_dbfs": round(db(rms), 1), "peak_dbfs": round(db(pk), 1),
"present": bool(db(rms) > PRESENCE_DBFS), "dur_s": round(len(mono) / sr, 1),
"bands": band_shares(mono, sr)}
row["hf_share_4k_16k"] = round(row["bands"]["4000-8000"]
+ row["bands"]["8000-16000"], 2)
return row
def compare_global(sect: Path, glob_path: Path, offset_s: float, dur_s: float) -> dict | None:
"""Same span, two separations. Correlation says whether the model made the
same decision; the RMS delta says whether normalisation changed the level."""
if not (sect.exists() and glob_path.exists()):
return None
a, sr = read(sect)
with sf.SoundFile(str(glob_path)) as f:
start = int(offset_s * f.samplerate)
if start >= f.frames:
return None
f.seek(start)
b = f.read(int(dur_s * f.samplerate), dtype="float32", always_2d=True).mean(axis=1)
k = min(len(a), len(b))
if k < sr:
return None
a, b = a[:k], b[:k]
denom = float(np.sqrt((a ** 2).sum() * (b ** 2).sum()))
corr = float((a * b).sum() / denom) if denom > 0 else 0.0
ra, rb = np.sqrt((a ** 2).mean()), np.sqrt((b ** 2).mean())
return {"corr": round(corr, 4),
"rms_delta_db": round(db(rb) - db(ra), 2),
"compared_s": round(k / sr, 1)}
def slug(title: str) -> str:
keep = "".join(c if c.isalnum() or c in " -_" else "" for c in title)
return "_".join(keep.split()).lower()[:48]
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("spec")
ap.add_argument("--out", help="JSON report path (default: releaseRoot/eda_stems.json)")
ap.add_argument("--no-global", action="store_true")
a = ap.parse_args()
spec = json.loads(Path(a.spec).read_text())
root = Path(spec["releaseRoot"])
stems = root / "stems_demucs"
segs = json.loads(Path(spec["segmentsRelease"]).read_text())
out = Path(a.out) if a.out else root / "eda_stems.json"
report, missing = [], []
for s in segs:
name = f"{s['track']:02d}-{slug(s['title'])}"
d = stems / "sections" / name
row = {"track": s["track"], "title": s["title"], "start": s["start"],
"duration": s["duration"], "dir": name, "stems": {}}
for src in SOURCES:
r = stem_row(d / f"{src}.wav")
if r is None:
missing.append(f"{name}/{src}")
continue
if src == "drums":
if r["present"]:
mono, sr = read(d / f"{src}.wav")
r["kick"] = kick_activity(mono, sr)
else:
# A kick percentage computed on an empty stem measures the
# noise floor and reads as a real number. Refuse it.
r["kick"] = {"active_pct": None, "reason":
f"drums stem absent ({r['rms_dbfs']} dBFS RMS)"}
if not a.no_global:
r["vs_global"] = compare_global(d / f"{src}.wav",
stems / "global" / f"{src}.wav",
s["start"], s["duration"])
row["stems"][src] = r
report.append(row)
# ---- printed summary: the three questions, in order --------------------
print(f"{'#':>3} {'title':<30} " + " ".join(f"{s[:6]:>7}" for s in SOURCES)
+ f" {'kick%':>6} {'HF4-16k':>8}")
print("-" * 96)
for r in report:
st = r["stems"]
cells = []
for src in SOURCES:
v = st.get(src)
cells.append(f"{v['rms_dbfs']:>7.1f}" if v else f"{'—':>7}")
kick = st.get("drums", {}).get("kick", {}).get("active_pct")
kick = kick if kick is not None else "n/a"
hf = st.get("drums", {}).get("hf_share_4k_16k")
print(f"{r['track']:>3} {r['title'][:30]:<30} " + " ".join(cells)
+ f" {kick if kick is not None else '—':>6} {hf if hf is not None else '—':>8}")
print("\n--- Q1 · is the 8-16 kHz deficit CYMBALS or CODEC NOISE? ---")
agg = {}
for src in SOURCES:
vals = [r["stems"][src]["bands"]["8000-16000"]
for r in report if src in r["stems"] and r["stems"][src]["present"]]
agg[src] = round(float(np.mean(vals)), 3) if vals else None
print(f" {src:<7} mean 8-16 kHz share of its own energy: {agg[src]}%")
known = {k: v for k, v in agg.items() if v is not None}
if known:
top = max(known, key=known.get)
others = [v for k, v in known.items() if k != top]
runner = max(others) if others else 0.0
# Ratio against the RUNNER-UP, not the minimum: one stem legitimately
# measuring 0.000% would otherwise divide by the epsilon guard and print
# a six-figure "spread" that means nothing.
spread = known[top] / runner if runner > 0 else float("inf")
print(f" -> highest in '{top}' ({known[top]}%), runner-up {runner}% "
+ (f"-> {spread:.1f}x" if spread != float("inf") else "-> runner-up is zero"))
print(" -> " + ("CYMBALS: the HF is drum content, so the air shelf can open up."
if top == "drums" and spread >= 2
else "NOT drum-dominated — treat the HF as noise, keep the shelf shut."))
print("\n--- Q2 · kick presence per track (40-120 Hz active time) ---")
ks = [(r["track"], r["title"], r["stems"]["drums"]["kick"]["active_pct"])
for r in report if "drums" in r["stems"] and "kick" in r["stems"]["drums"]
and r["stems"]["drums"]["kick"].get("active_pct") is not None]
if ks:
print(f" mean {np.mean([k[2] for k in ks]):.1f}% · "
f"worst: " + ", ".join(f"#{t} {ti[:18]} {p}%"
for t, ti, p in sorted(ks, key=lambda x: x[2])[:3]))
if not a.no_global:
print("\n--- Q3 · local vs global separation (same span, two normalisations) ---")
for src in SOURCES:
cs = [r["stems"][src]["vs_global"] for r in report
if src in r["stems"] and r["stems"][src].get("vs_global")]
if cs:
print(f" {src:<7} corr {np.mean([c['corr'] for c in cs]):.4f} "
f"· level delta {np.mean([c['rms_delta_db'] for c in cs]):+.2f} dB "
f"(n={len(cs)})")
if missing:
print(f"\n⚠ {len(missing)} stem file(s) missing — job may still be running")
Path(out).write_text(json.dumps({"gig": spec["gig"], "sections": report,
"hf_by_stem": agg, "missing": missing},
indent=1) + "\n")
print(f"\n✓ {out}")
return 0
if __name__ == "__main__":
sys.exit(main())
#!/usr/bin/env bash
# finish_cosmic — everything downstream of the masters, unattended.
#
# PLN went to bed with "your autonomous now going to bed tell me results tomo".
# The master and demucs runs are systemd units so they outlive the session, but
# the steps AFTER them were sitting in my session — which means a teardown would
# have left him with rendered masters and nothing else. So the rest is a unit
# too: it waits for the masters, splits the album, runs the stem EDA, and leaves
# a summary on disk that stands on its own.
#
# Ordered after cosmic-master.service in the unit file, so systemd does the
# waiting rather than a polling loop.
set -uo pipefail
TT=/home/pln/Work/Sound/Tidal/armada/tide-table
SPEC=$TT/judge_specs/cosmicfest-2026.json
ROOT=/home/pln/Work/Sound/Prod/Cosmic26_master
PY=/home/pln/Work/Sound/tidal-ears/.venv/bin/python
SUM=$ROOT/FINISH_SUMMARY.txt
exec > >(tee -a "$SUM") 2>&1
echo "=============================================================="
echo "finish_cosmic $(date -Is)"
echo "=============================================================="
if [ ! -f "$ROOT/Cosmic26_v1_streaming.flac" ]; then
echo "FATAL: no streaming master — the master unit did not produce one."
exit 1
fi
echo
echo "--- masters on disk ---"
for f in "$ROOT"/Cosmic26_v1_*.flac; do
[ -e "$f" ] || continue
printf '%s %s ' "$(basename "$f")" "$(du -h "$f" | cut -f1)"
ffprobe -v error -show_entries format=duration -of csv=p=0 "$f"
done
# The spec now lists only the streaming variant, so this splits exactly what was
# rendered. The -9 club target moved to `variants_deferred` — not a failure, a
# decision recorded there with its reasoning, and moving the key back renders it.
echo
echo "--- split + verify ---"
"$PY" "$TT/render_release.py" "$SPEC" --only split
"$PY" "$TT/render_release.py" "$SPEC" --only verify
echo
echo "--- stem EDA ---"
# Gate on LENGTH, not existence. A soundfile writer creates the file the instant
# it opens, so `-f` is true from the first second of an hour-long pass — the same
# trap that made me misread "4 of 14 sections done" earlier tonight. The global
# comparison seeks into these by absolute time, so a truncated global stem would
# silently mis-align every later track rather than fail.
# Compare against the MASTER's own duration, not a hardcoded number. A threshold
# of 3700 was a rubber stamp: the pass writes in 120 s chunks, so it sits at
# 3720 s — 31 chunks, one short — for minutes while looking "basically done" to
# any loose comparison. My own monitor fired a premature COMPLETE on that number.
WANT=$(ffprobe -v error -show_entries format=duration -of csv=p=0 \
"$ROOT/ZOOM0067.MP3" 2>/dev/null | cut -d. -f1); WANT=${WANT:-3779}
GDUR=0
if [ -f "$ROOT/stems_demucs/global/vocals.wav" ]; then
GDUR=$(ffprobe -v error -show_entries format=duration -of csv=p=0 "$ROOT/stems_demucs/global/vocals.wav" 2>/dev/null | cut -d. -f1)
GDUR=${GDUR:-0}
fi
if [ "$GDUR" -ge "$((WANT - 2))" ]; then
echo "global stems complete (${GDUR}s of ${WANT}s) — full EDA with local-vs-global"
"$PY" "$TT/eda_stems.py" "$SPEC"
else
echo "global stems SHORT (${GDUR}s of ${WANT}s) — section-only EDA."
echo " A short global file does not FAIL the comparison — compare_global takes"
echo " min(len(a),len(b)) — it compares fewer seconds per track and reports"
echo " thinner evidence under the same headings. Worth saying out loud."
"$PY" "$TT/eda_stems.py" "$SPEC" --no-global
fi
echo
echo "--- the air-shelf A/B, for PLN's ears ---"
# The EDA proved the 8-16 kHz deficit is cymbals (drums 0.367% vs 0.084%
# runner-up, 4.4x), which licenses opening the shelf but cannot say how far.
for pair in "+5:$ROOT/Cosmic26_v1_streaming.flac" "+8:$ROOT/ab_air8/streaming.flac"; do
g=${pair%%:*}; f=${pair#*:}
if [ -f "$f" ]; then
printf 'air %s dB %s ' "$g" "$(du -h "$f" | cut -f1)"
ffprobe -v error -show_entries stream=sample_rate,channels \
-show_entries format=duration -of csv=p=0 "$f" | tr '\n' ' '
echo
else
echo "air $g dB MISSING ($f)"
fi
done
# Numbers beside the A/B, because "which do you prefer" is an easier question
# when you also know what differs. Measured on the RENDERED files: a shelf set to
# +3 dB does not necessarily put +3 dB in the master, since the limiter
# downstream has opinions.
if [ -f "$ROOT/Cosmic26_v1_streaming.flac" ] && [ -f "$ROOT/ab_air8/streaming.flac" ]; then
echo
echo "--- A/B measured side by side ---"
"$PY" "$TT/compare_masters.py" \
"$ROOT/Cosmic26_v1_streaming.flac" "$ROOT/ab_air8/streaming.flac" \
--label "air +5" --label "air +8"
fi
echo
echo "--- what is on disk now ---"
for v in streaming; do
d="$ROOT/tracks_v1_$v"
[ -d "$d" ] && echo "$v: $(ls "$d"/*.flac 2>/dev/null | wc -l) tracks"
done
echo
echo "finish_cosmic done $(date -Is)"
{
"_comment": [
"Gig spec for COSMICFEST 2026-08-22 — jour J of cosmicfest_v2.67Hz. The date is",
"sourced in cosmicfest-2026_gigmeta.json and CORRECTS an unsourced 2026-08-23",
"this file carried until 2026-09-02. Per POSTPROD.md a new gig is a copy of",
"this file, never a copy of a script — this one is the first STEMLESS gig, so",
"read the differences from opal26.json as the point of the file:",
"",
" * NO stems. One stage-mic recording, so `stemsDir`/`stemmap`/`clipsDir`/",
" `keeps` are absent, not empty. Verified: apply_boundaries and",
" render_release read none of them; they need `segments`, `master`,",
" `variants`, `releaseRoot`, `releaseTag` only. build_judge_set DOES need",
" stems, which is exactly why the nominal segments came from",
" setlist_to_segments.py instead.",
" * NO premix is possible. OPAL's v4 was a per-orbit pass (it trimmed d4,",
" which owned 76-91% of the LF bed). Nothing here has an analogue: any",
" fix is broadband, or comes from demucs-separated stems used as a lens.",
" * `master` is the RAW recording because that is the timeline PLN's ear",
" calls were made against. `variants` point at mastered files that DO NOT",
" EXIST YET — the mix pass is the next step, and render_release will say",
" MISSING until it runs.",
" * `album` is deliberately ABSENT. CosmicFest 2026 has no page under",
" content/lives/2026/, so there is no canonical set title to copy;",
" inventing one would violate feedback_metadata_vs_mastering. Ask PLN."
],
"gig": "cosmicfest-2026",
"date": "2026-08-22",
"artist": "ParVagues",
"liveRoot": "/home/pln/Work/Sound/Tidal",
"source": "/home/pln/Work/Sound/Prod/Cosmic26_master/ZOOM0067.MP3",
"sourceProvenance": {
"device": "stage microphone, Zoom recorder",
"format": "MP3 320 kbps CBR, 44100 Hz, stereo",
"duration_s": 3779.318,
"sha256": "7a19758994dbfd121b1a30b010ac0be272a4f95a327ba666147ab548fe7d671b",
"measured_2026-09-01": {
"integrated_lufs": -25.2,
"true_peak_dbfs": -5.1,
"lra_lu": 11.1,
"tool": "ffmpeg ebur128=peak=true, full duration",
"reading": "Not clipped and not crushed — conservatively recorded, simply quiet. ~+11 dB to reach -14 LUFS, which needs limiting but is a gain problem, not a repair job.",
"caveat": "Lossy source AND a live FOH chain that people adjusted mid-set. Every derived master inherits both. PLN: 'we can only do our best, 80/20 likely'."
},
"copies": [
"/home/pln/Work/Sound/Prod/Cosmic26_master/ZOOM0067.MP3",
"/mnt/freebox/PLN/Work/Sound/Prod/Cosmic26_master/ZOOM0067.MP3"
],
"note": "Was a single copy in ~/Downloads until 2026-09-01. Do not confuse with Prod/cosmicfest/ or Prod/cosmicfestv0.live1.* — those are the 2025 edition (files dated 2025-06-27/28)."
},
"master": "/home/pln/Work/Sound/Prod/Cosmic26_master/ZOOM0067.MP3",
"segments": "/home/pln/Work/Sound/Prod/Cosmic26_master/segments_nominal.json",
"segmentsRelease": "/home/pln/Work/Sound/Prod/Cosmic26_master/segments_v1.json",
"releaseRoot": "/home/pln/Work/Sound/Prod/Cosmic26_master",
"releaseTag": "v1",
"variants": {
"streaming": "/home/pln/Work/Sound/Prod/Cosmic26_master/Cosmic26_v1_streaming.flac"
},
"earBoundaries": "/home/pln/Work/Sound/Tidal/armada/tide-table/judge_specs/cosmicfest-2026_ear_verified.json",
"earSetlist": "/home/pln/Work/Sound/Tidal/armada/tide-table/judge_specs/cosmicfest-2026_setlist_ear.json",
"titles": "/home/pln/Work/Sound/Tidal/armada/tide-table/judge_specs/cosmicfest-2026_titles.json",
"binS": 2.0,
"note": "14 tracks. 10 starts are PLN playhead calls; #1 is the origin; #12 #13 #14 were INFERRED from his notes on skipped boundaries and have never been heard as cuts. Those three are the ones to audition before any render is called final.",
"variants_deferred": {
"club": "/home/pln/Work/Sound/Prod/Cosmic26_master/Cosmic26_v1_club.flac",
"_why": [
"NOT rendered, deliberately, and not a failure. Two axes were being",
"conflated: PLN's 'one live club mix' means the CONTINUOUS set as opposed",
"to the album split — a FORM. This `club` entry is a -9 LUFS LOUDNESS",
"target, our own convention. Since no CosmicFest track is cut, the",
"continuous mix simply IS the master file, so both of PLN's forms are",
"already covered by the streaming master plus tracks_v1_streaming/.",
"It would also have missed: reaching -9 from a -25.6 LUFS source with",
"4.7 dB of linear headroom tripped the LRA floor at -10.5 on the trial,",
"i.e. the target costs more dynamics than it is worth on this source.",
"The CPU went to an air-shelf A/B instead, which answers a question PLN",
"actually has. Move this key back into `variants` to render it."
]
},
"abVariants": {
"_why": [
"The stem EDA settled that the 8-16 kHz deficit is CYMBALS, not codec",
"noise: drums hold 0.367% of their own energy above 8 kHz against 0.084%",
"for the runner-up, a 4.4x concentration. That licenses opening the air",
"shelf but cannot say how far, which is taste — so two full masters are",
"rendered, identical but for the shelf, and judged in place rather than",
"on a clip (OPAL's reverb lesson)."
],
"air+5": "/home/pln/Work/Sound/Prod/Cosmic26_master/Cosmic26_v1_streaming.flac",
"air+8": "/home/pln/Work/Sound/Prod/Cosmic26_master/ab_air8/streaming.flac"
}
}
{
"gig": "cosmicfest-2026",
"_provenance": [
"Generated by setlist_to_segments.py from cosmicfest-2026_setlist_ear.json",
"and cosmicfest-2026_boundaries_ear.json. DO NOT HAND-EDIT: edit those two and re-run.",
"`verified` holds ONLY starts PLN called on the playhead. Starts absent",
"here were inferred from his notes on SKIPPED boundaries and stay",
"`nominal` in apply_boundaries' output, so the two grades never merge."
],
"verified": {
"2": {
"start": 406.63,
"source": "playhead",
"boundary": 1,
"note": "from here its clearly something about drums, might even start earlier idk"
},
"3": {
"start": 698.63,
"source": "playhead",
"boundary": 3,
"note": "crossfade starts a bit earlier, this 11:38.6 is a good start point"
},
"4": {
"start": 1095.32,
"source": "playhead",
"boundary": 5,
"note": "that's the exact first sound of the next track you_my_sunshine"
},
"5": {
"start": 1364.91,
"source": "playhead",
"boundary": 6,
"note": "at 22:42 its end of reose_rouge, theres a loud blip noise, at 22:44.9 its the start of rose rouge, could we cut noise and blend better?"
},
"6": {
"start": 1897.76,
"source": "playhead",
"boundary": 7,
"note": "at 31:37.8 we're enough in the crosfade it makes a good 5 drops start."
},
"7": {
"start": 2145.61,
"source": "playhead",
"boundary": 8,
"note": "starting at 35:45.6 i hear proper only piment start sound"
},
"8": {
"start": 2525.44,
"source": "playhead",
"boundary": 11,
"note": "start of ouais_je_funk for real"
},
"9": {
"start": 2735.79,
"source": "playhead",
"boundary": 12,
"note": "start of perfect! there was no gimme_acid in this performance."
},
"10": {
"start": 2907.36,
"source": "playhead",
"boundary": 13,
"note": "this is end of perfect -> perfect cut at 48:27.4 into PunkAChien!"
},
"11": {
"start": 3072.47,
"source": "playhead",
"boundary": 14,
"note": "proper start of mafia"
}
},
"edits": {},
"cut_from_release": {}
}
{
"_what": [
"The fields of CosmicFest-2026's canonical www page that are NOT derivable from",
"the audio, the scores or the segments. Consumed by build_gig_tracksjson.py --meta.",
"Every entry carries {value, source, locator, date} per feedback_metadata_provenance,",
"so a reader in six months can tell PLN's words from a page-derived fact from a",
"drafted line. Nothing here is a guess: the two fields nobody could source are",
"absent, not filled in."
],
"title": {
"value": "CosmicSet 2026",
"source": "PLN",
"locator": "chat 2026-09-02: \"title id go with 'CosmicSet 2026'?\"",
"date": "2026-09-02"
},
"venue": {
"value": "CosmicFest",
"source": "PLN",
"locator": "chat 2026-09-02: \"'CosmicFest' is the venue\"",
"date": "2026-09-02"
},
"date": {
"value": "2026-08-22",
"source": "PLN, corroborating Web/www — and it CORRECTS the unsourced 2026-08-23 this gig's spec carried until 2026-09-02",
"locator": "PLN chat 2026-09-02: \"yea it was samedi 22 indeed :)\" · www/PRODUCT.md:21 \"cosmicfest_v2.67Hz (20-23 August 2026, jour J Saturday 22)\" + www/app/cosmicfest/CosmicFestClient.js:234 \"le 22 aout, dans le jardin\" (the lineup section) + :396 \"jour J - samedi 22 aout 2026\"",
"date": "2026-09-02"
},
"stage": {
"value": "Le Jardin",
"source": "derived from the invite page's lineup heading — PLN to confirm the name and casing",
"locator": "www/app/cosmicfest/CosmicFestClient.js:234 \"le 22 aout, dans le jardin\"",
"date": "2026-09-02"
},
"location": {
"value": "Labenne-Océan, France",
"source": "Web/www — canonical",
"locator": "www/PRODUCT.md:14 \"Labenne-Ocean (French Atlantic coast)\" + www/app/cosmicfest/page.js:4",
"date": "2026-09-02"
},
"description": {
"value": "CosmicSet 2026 — live coding set @ cosmicfest_v2.67Hz, Labenne-Océan, dans le jardin",
"source": "drafted from the canonical fields above — no fact of its own; PLN to approve the wording",
"locator": "mirrors the shape of content/lives/2026/opal-festival-2026.md's description",
"date": "2026-09-02"
},
"ctaURL": {
"value": "https://me.nech.pl/cosmicfest",
"source": "the edition's own invite micro-site, same CTA the 2025 page uses",
"locator": "www/content/lives/2025/cosmicfest.md ctaURL + the live /cosmicfest App Router route",
"date": "2026-09-02"
},
"ctaText": {
"value": "cosmicfest_v2.67Hz",
"source": "the festival's own wordmark, verbatim casing",
"locator": "www/DESIGN.md:53 \"Wordmark: cosmicfest (display) + v2.67Hz (mono badge)\"",
"date": "2026-09-02"
},
"tags": {
"value": [
"livecoding",
"cosmicfest",
"labenne",
"tidalcycles",
"summer",
"live"
],
"source": "same tag vocabulary as the 2025 CosmicFest page",
"locator": "www/content/lives/2025/cosmicfest.md tags",
"date": "2026-09-02"
},
"sections": {
"1": "Ouverture",
"2": "SUNSET",
"3": "SUNSET",
"4": "We call it NuJazz",
"5": "We call it NuJazz",
"6": "We call it NuJazz",
"7": "We call it NuJazz",
"8": "We call it NuJazz",
"9": "We call it NuJazz",
"10": "NUIT",
"11": "NUIT",
"12": "NUIT",
"13": "FINALE",
"14": "Outro"
},
"_sections_provenance": {
"source": "PLN's own set plan in backlog.md — the canonical informal tracklist source (reference_gig_tracklist_sources). NOT invented.",
"locator": "Tidal/backlog.md ## COSMICFEST (~line 1986): Ouverture / SUNSET / _We call it NuJazz_ / NUIT / FINALE",
"date": "2026-09-02",
"caveat": "The planned order matches the measured set exactly, with ONE difference: the plan lists 'WAP' [133] second in Ouverture and it does not appear in the recording, so the mapping shifts up by one from track 2 onward. Track 14 ('Outro: La Dub Sirene') is not in the plan at all - its section is taken from PLN's own title for it.",
"open_question": "Was WAP played? Track 1 is 6.8 min, in line with other single tracks (3 is 6.6, 7 is 6.3), so it does not look like two tracks merged - but that is an inference, not a measurement."
}
}
{
"gig": "cosmicfest-2026",
"_provenance": [
"Release titles PLN could not be derived from: tracks absent from the OPAL-26",
"release (the ear-signed title source) AND from catalog.generated.json.",
"Given verbatim by PLN in conversation, 2026-09-01. Per feedback_metadata_provenance",
"each entry carries its source; per feedback_metadata_vs_mastering these are",
"RELEASE metadata and belong to the gig, not to the mastering scripts.",
"`score` resolves a gig slug whose name does not match its .tidal filename."
],
"titles": {
"rose_rouge": {
"title": "Rose Rouge",
"source": "user",
"locator": "PLN, conversation 2026-09-01",
"as_of": "2026-09-01"
},
"mafia": {
"title": "Mafia sans Serif",
"score": "live/collab/raph/mafia_sans_serif.tidal",
"source": "user",
"locator": "PLN, conversation 2026-09-01",
"as_of": "2026-09-01",
"note": "The OPAL-26 release shipped this same score as plain \"Mafia\"; PLN's CosmicFest call is the fuller name. The gig slug `mafia` does not match the filename, hence `score`."
},
"livecode_parade": {
"title": "LiveCode Parade",
"source": "user",
"locator": "PLN, conversation 2026-09-01",
"as_of": "2026-09-01"
},
"Outro Dub Siren": {
"title": "Outro: La Dub Sirène",
"source": "user",
"locator": "PLN, conversation 2026-09-01",
"as_of": "2026-09-01",
"note": "PLN gave this lowercase (\"outro: la dub sirene\") and asked for proper French accents; the accent on sirène is his instruction, the capitalisation is a house-style guess. No .tidal exists — this is a played outro, not a composed track, so the catalog (keyed by score path) can never hold its title."
}
}
}
#!/usr/bin/env python3
"""master_stemless — master a gig captured on ONE file (no stems).
MASTERING.md's canonical chain assumes 12 Ardour orbit-stems: gain-stage each,
sum, spatialise, glue, push. A stage-mic capture has no stems, so the per-stem
half is simply unavailable and everything must happen on the bus. This is that
chain with the stem stages removed and ONE stage added, because the measurement
demanded it.
Measured on CosmicFest 2026 (ZOOM0067.MP3, full 62:59) before writing any filter:
L/R correlation 0.833 genuinely stereo — bass-mono is safe
<30 Hz 0.07% HPF is free headroom, not a compromise
300-1000 Hz 34.75% boxy midrange DOMINATES
1-4 kHz 14.76%
4-8 kHz 1.75% -17.6 presence nearly absent
8-16 kHz 0.12% -29.4 almost no air at all
I -25.2 LUFS · TP -5.1 dBFS · LRA 11.1 LU
That is the signature of a microphone in a room rather than a desk feed: the
PA's treble never reaches the mic position, the room absorbs what does, and
mid-bass builds up. So the dominant move is a broadband TILT, not gain — which
is why this file exists instead of a `--target-lufs` flag on the stem mixer.
Restraint where the measurement cannot support enthusiasm: the air shelf defaults
to +5, not the +12 the raw deficit suggests, because a 320 kbps MP3 puts codec
residue in the 8-16 kHz band alongside cymbals and a big boost lifts hiss into the
master.
**The demucs stem EDA has since settled which it is: CYMBALS.** `drums` holds
0.367% of its own energy above 8 kHz against 0.084% for the runner-up — a 4.4x
concentration, where evenly-spread energy would have meant noise. So opening the
shelf is licensed by measurement. HOW FAR is not: that is taste, so `--air-gain`
exists and the answer comes from an A/B of two full masters rather than from a
number I liked (`feedback_trim_threshold_ear_correction`, and OPAL's lesson that
an effect judged on a clip is not the same effect judged in place).
Bass-mono is done in MID/SIDE, not by splitting and re-summing bands. A
`lowpass(120)` summed with a `highpass(120)` puts a phase notch at the crossover;
high-passing the SIDE channel alone removes low-frequency stereo without the mid
path ever being filtered.
"""
from __future__ import annotations
import argparse
import json
import re
import subprocess
import sys
from pathlib import Path
# --- the chain, as data so the render can print exactly what it applied -------
HPF = 25 # measured: <30 Hz is 0.07% of total energy
LPF = 19500 # a 320k MP3 has nothing above ~20 k; strips codec junk
BOX_F, BOX_W, BOX_G = 450, 1.1, -2.5 # the 34.75% pileup at 300-1000
PRES_F, PRES_W, PRES_G = 3500, 0.9, 2.0 # articulation, gently
AIR_F, AIR_G = 6500, 5.0 # the 1.75% / 0.12% deficit, restrained
BASS_MONO_F = 120
# A loudness target reached by flattening the music is not reached. Below this
# residual range the push is doing harm, so we stop and report the miss.
LRA_FLOOR = 1.5
GLUE = "acompressor=threshold=-22dB:ratio=1.5:attack=30:release=100:makeup=1"
TARGETS = { # MASTERING.md's table; TP -1 dBTP for every platform
"streaming": dict(I=-14, TP=-1.0, LRA=11),
"club": dict(I=-9, TP=-1.0, LRA=11),
}
def tone_chain(air_g: float = AIR_G) -> str:
"""Everything before loudness. One string so it is identical in both passes."""
return (
f"highpass=f={HPF},"
f"equalizer=f={BOX_F}:t=q:w={BOX_W}:g={BOX_G},"
f"equalizer=f={PRES_F}:t=q:w={PRES_W}:g={PRES_G},"
f"highshelf=f={AIR_F}:g={air_g},"
f"lowpass=f={LPF}"
)
def filter_complex(loudnorm: str, pre_gain: float | None = None,
out_rate: int = 44100, air_g: float = AIR_G) -> str:
"""Tone -> M/S bass-mono -> glue -> [staged push] -> loudnorm.
`pre_gain` is what makes the club target reachable, and getting it wrong is
instructive: with the limiter placed BEFORE any gain, the signal arriving at
it still peaked at -9.5 dBTP, so a -1 dBFS ceiling never engaged, and
`linear=true` — correctly refusing to exceed TP — capped the master at
-11.2 LUFS instead of -9. The gain has to come FIRST so the limiter has
something to catch. Measured, not reasoned: that failure is why this
parameter exists.
Then MASTERING.md's staged advice, because 16 dB is too much for one stage:
a compressor takes the top off the loud passages, and the limiter only
handles what survives. One limiter doing all 16 dB pumps audibly.
"""
parts = [
f"[0:a]{tone_chain(air_g)},asplit=2[a][b]",
"[a]pan=mono|c0=0.5*c0+0.5*c1[M]",
f"[b]pan=mono|c0=0.5*c0-0.5*c1,highpass=f={BASS_MONO_F}[S]",
"[M][S]join=inputs=2:channel_layout=stereo[ms]",
# back to L/R: L = M+S, R = M-S
"[ms]pan=stereo|c0=c0+c1|c1=c0-c1[lr]",
f"[lr]{GLUE}[g]",
]
last = "g"
if pre_gain is not None:
parts.append(
f"[{last}]volume={pre_gain:.2f}dB,"
"acompressor=threshold=-14dB:ratio=2:attack=10:release=150:makeup=1,"
"alimiter=limit=0.891:attack=5:release=100:level=disabled[push]")
last = "push"
# ffmpeg's loudnorm declares its output at 192 kHz, because true-peak
# detection oversamples 4x and the filter does not resample back. Left alone
# it silently produces a 192 kHz master: the first CosmicFest streaming
# render came out correct on loudness (I -14.0, LRA 7.5, peak -1.0) and
# 1.26 GB at 192 kHz from a 44.1 kHz source — 4.3x the size it should be,
# at a rate no platform wants. topic_postprod_mastering already lists this
# under "loudnorm traps"; I hit it anyway by not re-reading it.
#
# So every chain ends by resampling back to the SOURCE rate. soxr because
# the default resampler is not worth the pennies saved on a mastering pass.
parts.append(f"[{last}]{loudnorm},aresample={out_rate}:resampler=soxr[out]")
return ";".join(parts)
def measure(src: Path, pre_gain: float | None = None,
out_rate: int = 44100, air_g: float = AIR_G) -> dict:
"""loudnorm pass 1. Two-pass is not optional: pass 2 needs measured_* to
apply ONE fixed offset (linear=true). A single-pass loudnorm rides the gain
and smears transients — the exact thing -14 LUFS is chosen to preserve."""
fc = filter_complex("loudnorm=print_format=json", pre_gain, out_rate, air_g)
p = subprocess.run(["ffmpeg", "-nostdin", "-hide_banner", "-i", str(src),
"-filter_complex", fc, "-map", "[out]", "-f", "null", "-"],
capture_output=True, text=True)
blocks = re.findall(r"\{[^{}]*\"input_i\"[^{}]*\}", p.stderr, re.S)
if not blocks:
print(p.stderr[-2500:], file=sys.stderr)
raise SystemExit("loudnorm pass 1 produced no measurement")
return json.loads(blocks[-1])
def render(src: Path, dest: Path, target: dict, m: dict,
pre_gain: float | None = None, out_rate: int = 44100,
air_g: float = AIR_G) -> None:
ln = (f"loudnorm=I={target['I']}:TP={target['TP']}:LRA={target['LRA']}:linear=true"
f":measured_I={m['input_i']}:measured_TP={m['input_tp']}"
f":measured_LRA={m['input_lra']}:measured_thresh={m['input_thresh']}"
f":offset={m['target_offset']}")
dest.parent.mkdir(parents=True, exist_ok=True)
cmd = ["ffmpeg", "-nostdin", "-hide_banner", "-v", "error", "-y", "-i", str(src),
"-filter_complex", filter_complex(ln, pre_gain, out_rate, air_g),
"-map", "[out]", "-ar", str(out_rate),
# 24-bit: the source is lossy, but the chain's gain and EQ produce
# values that deserve the headroom, and every downstream split
# re-encodes from this file.
"-c:a", "flac", "-sample_fmt", "s32", "-compression_level", "5", str(dest)]
r = subprocess.run(cmd, capture_output=True, text=True)
if r.returncode != 0:
print(r.stderr[-2500:], file=sys.stderr)
raise SystemExit(f"render failed: {dest}")
def verify(dest: Path) -> dict:
"""Re-measure the OUTPUT. feedback_verify_own_renders: the machine catches
objective errors so PLN's ears are spent on taste, not on arithmetic."""
p = subprocess.run(["ffmpeg", "-nostdin", "-hide_banner", "-i", str(dest),
"-af", "ebur128=peak=true", "-f", "null", "-"],
capture_output=True, text=True)
tail = p.stderr[-1800:]
def grab(label):
m = re.search(rf"{label}:\s*(-?[\d.]+)", tail)
return float(m.group(1)) if m else None
fmt = subprocess.run(
["ffprobe", "-v", "error", "-show_entries", "stream=sample_rate,channels",
"-of", "csv=p=0", str(dest)], capture_output=True, text=True).stdout.strip()
sr, ch = (fmt.split(",") + ["", ""])[:2]
return {"I": grab("I"), "LRA": grab("LRA"), "peak": grab("Peak"),
"sample_rate": int(sr) if sr.isdigit() else None,
"channels": int(ch) if ch.isdigit() else None}
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("spec")
ap.add_argument("--variant", action="append", choices=sorted(TARGETS),
help="default: every variant in the spec")
ap.add_argument("--source", help="override the spec's source (for testing on a slice)")
ap.add_argument("--out-dir", help="override releaseRoot")
ap.add_argument("--air-gain", type=float, default=AIR_G,
help=f"high-shelf gain at {AIR_F} Hz (default {AIR_G}). The stem EDA "
"settled that the 8-16 kHz deficit is CYMBALS (drums hold 0.367%% "
"of their energy there vs 0.084%% for the runner-up, 4.4x), so "
"opening this up is justified by measurement. HOW FAR is a taste "
"call — render an A/B and let PLN's ears decide, per "
"feedback_trim_threshold_ear_correction.")
a = ap.parse_args()
spec = json.loads(Path(a.spec).read_text())
src = Path(a.source or spec.get("source") or spec["master"])
variants = a.variant or list(spec["variants"])
root = Path(a.out_dir) if a.out_dir else Path(spec["releaseRoot"])
root.mkdir(parents=True, exist_ok=True)
out_rate = int(subprocess.run(
["ffprobe", "-v", "error", "-show_entries", "stream=sample_rate",
"-of", "csv=p=0", str(src)],
capture_output=True, text=True).stdout.strip().split("\n")[0])
print(f"source {src.name} ({out_rate} Hz — every render resamples back to this)")
print(f"tone {tone_chain(a.air_gain)}")
print(f" bass-mono (mid/side) <{BASS_MONO_F} Hz · glue 1.5:1 @ -22 dB\n")
report = {}
# One tone-only measurement serves every variant: the chain up to the push
# is identical, so measuring it per variant would burn a full pass each.
base = measure(src, None, out_rate, a.air_gain)
print(f"tone-only: I={base['input_i']} TP={base['input_tp']} LRA={base['input_lra']}\n")
for v in variants:
t = TARGETS[v]
headroom = t["TP"] - float(base["input_tp"]) # what a fixed offset can give
want = t["I"] - float(base["input_i"]) # what the target asks for
staged = want > headroom # a limiter is the only way there
pre_gain = want if staged else None
dest = root / f"{v}.flac" if a.out_dir else Path(spec["variants"][v])
print(f"=== {v}: I={t['I']} TP={t['TP']} · needs {want:+.1f} dB, "
f"linear headroom {headroom:+.1f} dB -> "
+ (f"STAGED (pre-gain {pre_gain:+.1f} dB, comp + limiter)" if staged
else "linear, no limiter"))
# Converge on the target. Every dB the limiter absorbs is a dB the final
# loudnorm cannot add without breaching TP, so the first pre-gain
# undershoots: the 90 s trial reached -10.5 against a -9 target. Feeding
# the residual back in converges, because the shortfall IS the gain the
# limiter ate. Bounded at 3 attempts, and guarded by a dynamics floor —
# `presence_is_a_precondition`: a target hit by flattening the music is
# not a target hit. If the floor trips we report the miss and keep the
# better-sounding render rather than clipping our way to a number.
attempts, got, m, best = [], None, base, None
for i in range(3):
m = measure(src, pre_gain, out_rate, a.air_gain) if staged else base
if staged:
print(f" attempt {i+1}: pre-gain {pre_gain:+.1f} dB -> "
f"I={m['input_i']} TP={m['input_tp']} LRA={m['input_lra']}")
render(src, dest, t, m, pre_gain, out_rate, a.air_gain)
got = verify(dest)
lra_out = got["LRA"] if got["LRA"] is not None else 0.0
hit = got["I"] is not None and abs(got["I"] - t["I"]) <= 0.5
attempts.append({"pre_gain_db": pre_gain, "out": dict(got)})
if best is None or (got["I"] is not None and best["I"] is not None
and abs(got["I"] - t["I"]) < abs(best["I"] - t["I"])):
best = dict(got)
if hit:
break
if lra_out < LRA_FLOOR:
print(f" stop: LRA {lra_out} is below the {LRA_FLOOR} LU floor — "
f"{t['I']} LUFS costs more than it is worth on this source")
break
if not staged or got["I"] is None:
break
pre_gain += (t["I"] - got["I"]) # the shortfall is the missing gain
ok = (got["I"] is not None and abs(got["I"] - t["I"]) <= 0.5
and got["peak"] is not None and got["peak"] <= t["TP"] + 0.3
# Loudness alone passed a 192 kHz master once. Format is spec too.
and got["sample_rate"] == out_rate and got["channels"] == 2)
squash = (float(base["input_lra"]) - got["LRA"]) if got["LRA"] is not None else None
print(f" out: I={got['I']} LRA={got['LRA']} peak={got['peak']} "
f"{got['sample_rate']}Hz/{got['channels']}ch "
f"{'✓' if ok else '✗ MISSED TARGET'}"
+ (f" · dynamics lost {squash:.1f} LU" if squash and squash > 1 else ""))
print(f" {dest}")
report[v] = {"target": t, "air_gain_db": a.air_gain, "measured_tone": base, "measured_in": m,
"measured_out": got, "in_spec": ok, "staged": staged,
"pre_gain_db": pre_gain, "lra_lost": squash,
"attempts": attempts, "path": str(dest)}
(root / "master_report.json").write_text(json.dumps(report, indent=1) + "\n")
print(f"\n✓ {root/'master_report.json'}")
return 0 if all(r["in_spec"] for r in report.values()) else 1
if __name__ == "__main__":
sys.exit(main())
#!/usr/bin/env python3
"""setlist_to_segments — build the nominal segments for a gig recorded WITHOUT stems.
`build_judge_set` derives nominal segments from per-orbit stem activity. A set
captured on a single stage mic has no stems, so that door is shut — but the ear
pass still happened, and `judge_specs/<gig>_setlist_ear.json` already holds one
row per track. This turns that file into the two inputs `apply_boundaries` wants:
1. a NOMINAL segments list (what the spec's `segments` key points at), and
2. an ear file in `apply_boundaries`' own schema — `verified` keyed by TRACK.
Why both, when the setlist already carries the ear's numbers: the setlist mixes
two grades of truth. Some starts are playhead calls PLN made while listening;
others were inferred from the notes on boundaries he SKIPPED. Splitting them
means `apply_boundaries` prints `ear` or `nominal` per row and asserts that
every playhead call reached the output — so an inferred edge can never be
mistaken later for one he actually heard.
Titles come from what SHIPPED, never from a slug: the OPAL-26 release is
ear-signed, so its `segments_v4.json` joined to that gig's `tracks.json` (score
path in performance order) is the authoritative slug -> release-title map. The
catalog is a second source. A track in neither is emitted as NEEDS_PLN rather
than title-cased into something plausible — `feedback_metadata_provenance`.
Declared BPM is parsed from the track's own `.tidal` (`setcps (N/60/4)`), which
is the score's claim, not a measurement. A tempo-knob range is recorded as a
range so nobody reads a floor as a fact.
"""
from __future__ import annotations
import argparse
import json
import re
import sys
from pathlib import Path
REPO = Path(__file__).resolve().parent.parent.parent
HERE = Path(__file__).resolve().parent
# `setcps (124/60/4)` — the BPM is the numerator. Also catches `cps (N/60/4)`.
CPS_STATIC = re.compile(r'setcps\s*\(\s*([0-9.]+)\s*/\s*60\s*/\s*4\s*\)')
# the knob idiom: `# cps ((range LO HI "^NN")/60/4)`
CPS_RANGE = re.compile(r'range\s+(\(?[0-9.+\- ]+\)?)\s+(\(?[0-9.+\- ]+\)?)\s+"\^\d+"')
def opal_title_map() -> dict[str, str]:
"""slug -> release title, from the one release that is signed off by ear.
The join key is performance order: `tracks.json` lists score paths 1..N in
the order played, and `segments_v4.json` carries `perf_track` alongside the
title that actually shipped on the file. Neither file alone has both halves.
"""
tj = Path("/home/pln/Work/Web/www/content/lives/2026/opal-festival-2026/tracks.json")
sv = Path("/home/pln/Work/Sound/Prod/Opal26_master/segments_v4.json")
if not (tj.exists() and sv.exists()):
return {}
scores = [t.get("file") or t.get("score") or "" for t in json.loads(tj.read_text())["tracks"]]
by_perf = {r["perf_track"]: r["title"] for r in json.loads(sv.read_text())}
out = {}
for i, path in enumerate(scores, start=1): # perf_track is 1-based
if path and i in by_perf:
out[Path(path).stem] = by_perf[i]
return out
def catalog_title_map() -> dict[str, str]:
p = HERE / "catalog.generated.json"
if not p.exists():
return {}
tracks = json.loads(p.read_text())["tracks"]
tracks = list(tracks.values()) if isinstance(tracks, dict) else tracks
return {Path(t["id"]).stem: t["title"] for t in tracks
if isinstance(t, dict) and t.get("id") and t.get("title")}
def find_score(slug: str) -> Path | None:
"""The score for a slug. One track = one standalone .tidal, per CLAUDE.md."""
hits = [p for p in REPO.rglob(f"{slug}.tidal") if ".git" not in p.parts and ".claude" not in p.parts]
return sorted(hits, key=lambda p: len(p.parts))[0] if hits else None
def declared_bpm(score: Path | None):
if not score or not score.exists():
return None
text = score.read_text(errors="replace")
live = [ln for ln in text.splitlines() if not ln.lstrip().startswith("--")]
body = "\n".join(live)
if (m := CPS_STATIC.search(body)):
return float(m.group(1))
if (m := CPS_RANGE.search(body)):
return {"range": [m.group(1).strip(), m.group(2).strip()], "note": "tempo is a knob"}
return None
def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--setlist", required=True, help="judge_specs/<gig>_setlist_ear.json")
ap.add_argument("--boundaries", required=True, help="judge_specs/<gig>_boundaries_ear.json")
ap.add_argument("--out-segments", required=True)
ap.add_argument("--out-ear", required=True)
ap.add_argument("--titles", help="judge_specs/<gig>_titles.json — PLN's authored "
"release titles for tracks no shipped release or catalog can name")
ap.add_argument("--tol", type=float, default=0.15,
help="seconds within which a setlist start counts as matching a playhead call")
a = ap.parse_args()
setlist = json.loads(Path(a.setlist).read_text())
bounds = json.loads(Path(a.boundaries).read_text())
decided = bounds.get("decided", {})
tracks = setlist["tracks"]
opal, cat = opal_title_map(), catalog_title_map()
authored = {}
if a.titles:
authored = json.loads(Path(a.titles).read_text()).get("titles", {})
segs, verified, needs_pln, unmatched = [], {}, [], []
for t in tracks:
n, slug, start = int(t["n"]), t["track"], float(t["start"])
# Which playhead call, if any, produced this start?
hit = next((k for k, v in decided.items()
if abs(float(v["start"]) - start) <= a.tol), None)
auth = authored.get(slug, {})
title = auth.get("title") or opal.get(slug) or cat.get(slug)
if not title:
title = f"NEEDS_PLN:{slug}"
needs_pln.append(slug)
# An authored `score` resolves a gig slug that is not its filename
# (CosmicFest's "mafia" is mafia_sans_serif.tidal).
score = (REPO / auth["score"]) if auth.get("score") else find_score(slug)
if score is None:
unmatched.append(slug)
segs.append({
"track": n,
"title": title,
"start": round(start, 3),
"bpm": declared_bpm(score),
"score": str(score.relative_to(REPO)) if score else None,
"slug": slug,
"title_source": ("authored" if auth.get("title") else
"opal-release" if slug in opal else
"catalog" if slug in cat else "NEEDS_PLN"),
"title_prov": ({k: auth[k] for k in ("source", "locator", "as_of") if k in auth}
or None),
"start_source": f"ear:boundary#{hit}" if hit else ("origin" if start == 0.0 else "inferred"),
})
if hit:
verified[str(n)] = {
"start": round(float(decided[hit]["start"]), 3),
"source": "playhead",
"boundary": int(hit),
"note": decided[hit].get("note", ""),
}
Path(a.out_segments).write_text(json.dumps(segs, indent=1) + "\n")
Path(a.out_ear).write_text(json.dumps({
"gig": setlist["gig"],
"_provenance": [
f"Generated by setlist_to_segments.py from {Path(a.setlist).name}",
f"and {Path(a.boundaries).name}. DO NOT HAND-EDIT: edit those two and re-run.",
"`verified` holds ONLY starts PLN called on the playhead. Starts absent",
"here were inferred from his notes on SKIPPED boundaries and stay",
"`nominal` in apply_boundaries' output, so the two grades never merge.",
],
"verified": verified,
"edits": {},
"cut_from_release": {},
}, indent=2) + "\n")
print(f"{'#':>3} {'start':>9} {'bpm':>7} {'start_source':<20} {'title_source':<13} title")
for s in segs:
bpm = s["bpm"]
bpm = f"{bpm:g}" if isinstance(bpm, (int, float)) else ("knob" if bpm else "—")
print(f"{s['track']:>3} {s['start']:>9.2f} {bpm:>7} "
f"{s['start_source']:<20} {s['title_source']:<13} {s['title']}")
print(f"\n{len(segs)} tracks · {len(verified)} playhead-called starts · "
f"{len(segs) - len(verified)} not")
if needs_pln:
print(f"⚠ {len(needs_pln)} title(s) NEED PLN (no shipped release, no catalog entry): "
+ ", ".join(needs_pln))
if unmatched:
print(f"⚠ {len(unmatched)} slug(s) with no .tidal found: " + ", ".join(unmatched))
print(f"✓ {a.out_segments}\n✓ {a.out_ear}")
return 0
if __name__ == "__main__":
sys.exit(main())
Markdown is supported
0% or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment