Commit c2fd9716 by PLN (Algolia)

docs(backlog): clear the tail, keep what was in it

PLN: "deliberate killed tail was just a agent log or srth, capture that removed
as learning, nut then let me clean that backlog indeed was gboing tedious."

The tail is cleared and the new AI ENGINEER candidate section stands. But the
543 lines were not a log: they were two sessions of rig notes from 6-7
September, and six items in them were still open. Losing them to a backlog
cleanup is the quiet kind of loss — nobody misses a note they have forgotten
writing.

So the whole block is archived verbatim in docs/, with the open items hoisted
to the top of it where they can be read without scrolling. The one that
matters: `sudo tools/install-protect.sh` was never run, so the OOM-protection
daemon has never actually protected anything — it reported ok only because
scsynth and sclang already sat at the target value, and no write was attempted.
parent 092e1a17
......@@ -2013,8 +2013,27 @@ ParVagues et Shipow
## AI ENGINEER :rocket:
120Mn X ~5mn = 24 candidatess
### CANDIDATES
- Quand on Decolle
### NUJAZZ PARADIZE
- Salut Nu
- Liquid Nite <3 TODO Finish convert
- Liquid Finale
- Love First
- **Take Five Drops**
- **Rose Rouge**
- L'Ete a Mauerpark [confirm d9 effects]
## Finale comedown
- [120] Sept1 TODO FIX TO NEW CONTROLLER
- [120] Cafe Glace
- [114] Revolution
......@@ -2038,620 +2057,3 @@ url: "/home/pln/Work/Art/GLITCHWAVE/outputs/bg_live4.gif"
url: "/home/pln/Work/Art/GLITCHWAVE/outputs/digital_flow.gif"
url: "/home/pln/Work/Art/GLITCHWAVE/outputs/sunrise_flow.gif"
url: "/home/pln/Work/Art/GLITCHWAVE/outputs/glitch_ocean_compressed.gif"
# Rig — opened 2026-09-06 (gig night on the new XPS24)
Floated during the session that got rose_rouge playing, not done. Full context:
`armada/tasks/completed-archive.md` (top entry) + `docs/2026-09-06-ardour-fresh-jig.md`.
## Visuals last mile — DONE 2026-09-06
Two corrections happened here, and both are worth more than the task was.
**First**, the plan in this slot ("one home `visuals/scenes/`, compress the
gifs, rewrite the refs repo-relative") proposed three things that were already
done or already decided against: the home exists and is the HUD's `scenesDir`
default (`package.json:101`); the gifs were compressed 2026-08-29
(`rose_bloom*.mp4`, 65-204 KB); and committing them was deliberately rejected
in `.gitignore:53-56` with the reason written in. A fence got kicked before
anyone read why it was there.
**Second**, the correction was itself half-wrong. "The refs work as designed,
nothing to rewrite" is true only on the box that authored the image.
`scenes.js:resolve()` tried an absolute spec and nothing else, and
`_preferDerived()` — the twin lookup — is only reached through a source that
EXISTS. So on the XPS24 every ref resolved to null with the mp4 sitting right
there, and the derived-twin mechanism was quietly conditional on being the
machine that made the gif.
Fixed in the HUD (`pulsar-parvagues-hud` 78b1755, **committed not pushed** —
that repo is on `main`): an absolute miss falls back to the stem against
scenesDir, the name `scene-ingest.cjs` already writes. SSOT design untouched;
it just stops requiring the SSOT to be local. Measured over the real corpus:
**0 of 5 distinct scene targets resolved before, 5 of 5 after**, and a target
with no twin still resolves to null.
Media synced: `visuals/scenes/`, 13 mp4s. Use `-H` — `decollage.mp4` and
`decolle.mp4` are hardlinks of one 120 MB file, so the 197 M source tree lands
as 124 M:
rsync -aH --partial xps22.local:'~/Work/Sound/Tidal/visuals/scenes/' visuals/scenes/
~~Still open, and a deliberate choice rather than an oversight~~ **SYNCED
2026-09-06.** `~/Work/Art/GLITCHWAVE` (3.1 G) is the SSOT *source* tree, needed
only to re-run `npm run scenes:ingest` for a NEW scene, never to play one. It
was deliberately left on xps22 on the grounds that authoring happens there —
PLN overruled that: "glitchwave cant hurt tbh". Fetched with `rsync -aH`, so
this box can now author scenes as well as play them, and the GIF `url:` lines
scattered through this file resolve locally.
rsync -aH xps22.local:'~/Work/Art/GLITCHWAVE/' ~/Work/Art/GLITCHWAVE/
## Sound / mix
- ~~**`Tidal 12` fader is at `-inf dB`**~~ **DONE, and the framing was wrong.**
PLN: "tidal 12 -inf is random i might turn on/off any fader its not signal
throughout perf it mioves always and i save random". Fader positions at save
time are arbitrary, so there was nothing to decide. Measured: **9 of 12**
faders had drifted, five to -inf. `gig-up.sh` now runs
`fader-baseline.py --restore` in the Ardour-closed window, full restore to the
2026-08-02 baseline. Nothing further unless the baseline itself should move
(`--capture`).
- ~~**`/usr/local/sbin/perf-audio` not deployed**~~ **DONE 2026-09-06** — installed
with its sudoers rule (`visudo -cf` parsed OK), so gig-up can set performance mode.
It had played at `gear: None, mode: silent` with the CPU capped at powersave.
- ~~**`parvagues-protect` not installed**~~ **DONE 2026-09-06, and the install
immediately exposed a bug in the guard itself** — fixed in `8f2c510`. It ran as
root and still could not write `/proc/<pid>/oom_score_adj`: the unit's
`CapabilityBoundingSet` omitted `CAP_DAC_OVERRIDE`, which is what uid 0 uses to
bypass file permissions, and the target is mode 0644 owned by `pln`. Every write
returned EACCES under an error message that said "need root, have uid 0".
scsynth and sclang showed `ok` only because they were already at the target, so
no write was attempted — the daemon had never protected anything.
**Still to verify:** reinstall and watch a restart re-protect on its own, which
is the only test that proves it (`systemctl --user restart parvagues-sc`, then
`parvagues-protect --check` — expect `oom:200->-1000` naming the NEW pid).
- 🔴 **`preload.scd` did not exist on the XPS24 — 60 banks would have been read
off disk mid-set.** Found 2026-09-07 while trying to measure the quantum, and
it is the same failure `tools/check-preload.sh`'s header calls out as having
"debuted as a crackle at the venue" — except on xps22 the plan existed and
warmed the wrong list, while here there was no plan at all.
`start_and_midi.scd:259` tests `File.exists("preload.scd")` and silently
takes the other branch: *"preload: no preload.scd — lazy-loading samples on
demand."*
**Measured, not inferred.** Driving `rose_rouge` headlessly for ~3 min logged
**46** `reading soundfile as needed` lines — `rose:28.0`, `rose:29.0`,
`rose:30.0`, one every 5-12 s, each one a disk read landing on the audio
thread. In the same window the SuperCollider node's `pw-top` **W/Q reached
1.300 at the rig's stock 1024** and 1.390 at 512, i.e. exceeding its deadline
regardless of quantum. `check-preload.sh` then said it plainly: *"60 bank(s)
would be read from DISK on first play (crackle, mid-transition)"*.
**Fixed by running `tools/check-preload.sh --fix`** — 60 banks, all 16
setlist tracks, 7.5 KB, gitignored (correctly, it is generated). Verified at
the next boot: `preload: warming the set's samples…` /
`=== PRELOAD: 60/61 banks OK in 2.6 s ===`.
- ~~**The preload plan is fresh and aimed at the wrong set.**~~ **FIXED
2026-09-07: `tools/gen_setlist.py`.** The preload now warms what PLN actually
PLAYS, computed from the canonical gig records rather than a hand-kept file,
so it cannot drift: **60 banks → 111**, covering **39 tracks** from the last
12 months instead of 16 from one August gig. 51 banks were one first-play
away from a disk read. `rose` — the bank behind all 41 measured lazy reads —
is covered.
Sources, in authority order: `<www>/content/lives/<year>/<slug>/tracks.json`
(canonical; carries both `date` and each track's exact repo-relative `file`)
for opal-festival-2026, montreuil-algorave, raise and bunker; then
`armada/tide-table/judge_specs/*_setlist_ear.json` for gigs whose tracks.json
is not built yet — that is where the **cosmicfest-2026** list lives, "THE
ground truth, 14 tracks", with `rose_rouge` at #5. Those carry no gig date, so
the tool includes them and SAYS SO rather than inventing one.
`check-preload.sh` now computes its own list instead of borrowing
`set-coherence.setlist_tracks()`, because the two ask different questions: a
cheat sheet wants TONIGHT's running order, a preload wants anything PLN might
play, and the cost asymmetry is total (a warmed unused bank costs boot
seconds; an unwarmed one costs a disk read on the audio thread at the venue).
Override with `PV_PRELOAD_SETLIST=path` for one specific gig, or
`PV_PRELOAD_MONTHS=6`.
**Measured cost of over-covering:** 124 banks eager, `122/124 OK in 5.0 s`
(was 60/61 in 2.6 s), scsynth RSS **4.8 G** — on a box with 62 G and 40 G
available, i.e. 8%. Fine. Two unresolved track names, both harmless:
`mafia` is an alias for `mafia_sans_serif.tidal`, already in via opal, and
`Outro Dub Siren` is a live improvisation with no file.
- ~~**AppleDouble resource forks masquerading as .wav**~~ **DELETED 2026-09-07
on PLN's explicit say-so** ("no reason to keep these imo"). **198 files, 0.3
MB**, every one confirmed by `file` magic before removal — the guard was
mechanical, not the name pattern, and 198/198 candidates came back
"AppleDouble encoded Macintosh file" with zero false positives.
69 were in playable banks: `rhadamanthe_fx` 227→189, `rhadamanthe_divers`
314→293, `rhadamanthe_vocal` 77→72, `rhadamanthe_melo` 32→30. The other 129
sat under `~/Work/Sound/Samples/baba/__MACOSX/`; the emptied dirs were
`rmdir`-ed (rmdir refuses a non-empty directory, so it cannot take data).
**Sample indices were NOT disturbed**: the `zz.` prefix sorted every one of
them after all real samples, which is presumably why someone renamed rather
than deleted. Only the counts moved. Boot is now free of the
`WARNING: File reading failed for path:` wall, and the preload reports
**124/124 banks OK in 5.2 s** where it used to report 122/124.
- ~~**`preload.scd`'s COUNT MISMATCH banner cries wolf**~~ **FIXED 2026-09-07,
and the real bug was underneath it.** Three separate faults, in order of
discovery:
1. The banner's headline said "whitelist is STALE, regenerate it" when the
cause was unreadable files, and regenerating provably changed nothing. It
now says "N bank(s) DID NOT FULLY LOAD", separates the two causes
(`got == 0` = whole bank failed; `0 < got < expected` = folder changed OR
some files are not audio) and prints the `file`-based one-liner that tells
them apart. *(Correction to an earlier claim of mine: it always DID name
the offending banks — `bad.do { ... }` prints them. My journal grep was
filtering those lines out. Only the headline was wrong.)*
2. **`check-preload.sh` compared bank NAME SETS, never counts.** So a bank
whose contents changed was invisible: the plan said `rhadamanthe_vocal, 77`
against a folder holding 72 and the checker printed **ok**, while
SuperDirt's boot said `expected 77 files, got 72`. Worse, `--fix` refused
to regenerate, because by the name-set test nothing had changed. The
everyday way to hit this is not deleting files — it is **dropping a new
sample pack into an existing bank**, which PLN does, and every new file
would then lazy-load mid-set under a green check. Now compares
`name count` pairs and reports `plan N -> disk M` per drifted bank.
3. **The bank-name regex was case-blind** — `[a-z0-9_]+` truncated `rampleA0`
to `rample` and collapsed every `vocalOoh1`/`vocalScatJ`/... into one
`vocal`. That is why the checker said **111 banks** while SuperDirt loaded
**124**: 124 was always the truth, and the shortfall was the regex, which
was wrong in both directions — undercounting the plan AND inventing two
bank names that do not exist. Found by asking why two numbers never
matched instead of assuming one rounded the other. Fixed to
`[A-Za-z0-9_]+`; and because `bank_counts()` inherited it, drift in any
camelCase bank had still been invisible after fault 2 was fixed.
**Now trustable, three ways agreeing**: `check-preload.sh` says 124, the boot
says `124/124 banks OK`, an independent Python parse says 124 names / 124
pairs / no gap. Drift detection tested in both directions on a camelCase bank
(`vocalOoh1` plan 9 vs disk 13 → STALE exit 1; restored → ok exit 0, plan
byte-identical after the test).
- (superseded) **64 AppleDouble resource forks are masquerading as .wav in
three banks PLN plays.** `rhadamanthe_fx` (38), `rhadamanthe_divers` (21),
`rhadamanthe_vocal` (5) — files named `zz._<original>.wav`, 4096 bytes each,
and `file` calls them exactly what they are: "AppleDouble encoded Macintosh
file". They are macOS metadata, not audio. SuperCollider fails to read them
(that is the `122/124` above and a wall of `WARNING: File reading failed for
path:` at every boot), and because they are counted as files in the bank
**they occupy sample indices** — so `s "rhadamanthe_fx:n"` is silent for 38
values of n, and `bank_file_count` overcounts. Another 129 sit under
`~/Work/Sound/Samples/baba/__MACOSX/`, which is extraction junk rather than a
playable bank.
**NOT touched**: the `zz.` prefix looks like a deliberate rename to sort them
last rather than delete them, which is a fence with a reason until PLN says
otherwise. Deleting them is his call; `find <bank> -name 'zz._*' -delete`
after a look is the whole fix, and it would also silence the boot warnings
and un-shift the indices.
- **`preload.scd`'s COUNT MISMATCH banner still cries wolf on a fresh plan**
(unchanged): `122/124 banks OK` then `=== PRELOAD COUNT MISMATCH — whitelist
is STALE, regenerate it ===` on a plan generated sixty seconds earlier. Now
we know WHY the counts differ — the 2 misses are the AppleDouble banks above,
which the generator counts and the server cannot load — so the banner is
reporting a real fact with the wrong diagnosis. It should say "N banks failed
to LOAD" and name them, not "regenerate the whitelist", which does nothing.
- 🔴 (superseded, kept for the shape) **The preload plan was fresh and aimed at
the wrong set.** Re-running the
same load after the fix still logged 41 lazy reads — and **all 41 were the
`rose` bank**, because `rose_rouge` is NOT in `armada/setlist_opal2026.txt`.
Zero of the 60 preloaded banks lazy-loaded, and SC logged ZERO `late`
messages in that window, so the mechanism is provably
working; it is warming the 16-track OPAL set while the track this box has
actually been playing (see the rig-state memory) is outside it.
This is `check-preload.sh`'s own documented failure one layer up: there, the
plan was stale because two tools disagreed about the setlist (10 vs 13
tracks). Here the plan matches its setlist exactly and **the setlist is stale
relative to what gets played**. A freshness check that compares the plan to
the setlist cannot see this — it needs to compare against what is actually
loaded. Cheapest honest fix: add current material to the setlist (or a
second "current" list the generator also reads), then regenerate.
- **Preload reports `60/61` and its own `PRELOAD COUNT MISMATCH — whitelist is
STALE, regenerate it`** on a plan generated sixty seconds earlier. So either
the generator counts one bank the server cannot load, or the two count
different things. Worth ten minutes: the mismatch banner is a real signal that
currently cries wolf on a fresh plan, and a warning that fires when nothing is
wrong is a warning nobody reads on the night it matters.
- ~~**Three global FX synthdefs are missing at every boot**~~ **FIXED 2026-09-07
with one `s.sync`.** `global_mi_verb2` / `global_mi_clouds2` /
`global_mi_ripples2` were reported `SynthDef not found` on every clean boot —
**84 error lines**, i.e. 3 defs × 14 orbits × 2 lines. So **every orbit's
MiVerb, MiClouds and MiRipples send was a dead node for the whole session**:
`# verbwet`, `# cloudswet`, `# ripplesreson` did nothing, silently, because a
missing global effect is an absence of effect rather than a noise you notice.
**mi-UGens was installed perfectly all along** — 11 `.so` + 11 classes at
`~/.local/share/SuperCollider/Extensions/mi-UGens` — and the SynthDefs were
defined correctly at `start_and_midi.scd:273-296`. The bug was ordering:
`SynthDef(...).add` is ASYNCHRONOUS (compiles locally, sends `/d_recv`,
returns) and the very next statement, `~dirt.orbits.do { ... x.initNodeTree }`,
sends `/s_new` for those names. The `/s_new` overtook the `/d_recv`. One
`s.sync` before the registration, legal because we are inside
`s.waitForBoot`'s Routine — the idiom was already sitting commented out at
line 236. Verified: 84 errors before, **0 after**, same file one line apart.
(`spectral-freeze` above does NOT need it: it replaces a def whose synths are
built per event, long after `/d_recv` lands.)
- **`rig-doctor` now monitors mi as gear** (2026-09-07, PLN's ask). Two checks
where there was one presence test:
* `mi-UGens extension` enumerates the 11 plugins **and** the 11 classes we
actually use, and WARNs naming whichever is absent — `d.is_dir()` is true
whether the directory holds a working install or a stale README.
* `SuperDirt boot errors` (new, under SERVICES) scopes the journal to the
CURRENT boot and FAILs on `SynthDef ... not found` / `FAILURE IN SERVER`,
naming the distinct SynthDefs rather than repeating one line per orbit.
The question it asks is generic, so it catches the next async-ordering bug
too, not just this one. WARNs (not FAILs) when parvagues-sc is down, since
on-demand is the correct state before a set.
Both tested in BOTH directions: the boot-error check FAILs on the recorded
pre-fix journal window (naming all three defs) and PASSes on the post-fix
boot. That test is also what caught `re` never having been imported in
`rig-doctor.py` — my new code was its first user, so the check would have
crashed the doctor the first time it found something.
Doctor now: **47 checks, 0 fail, 10 warn, 37 pass.**
**The mechanism was never broken**: `gig-up.sh:182` already calls the checker
and auto-fixes on a non-zero exit, and the checker does exit 1 correctly.
`gig-up.sh` simply has not been run on this box since the setlist grew. So
the gap is not the code, it is that the gig-up path is the only thing that
regenerates the plan and nothing else notices. **`rig-doctor` should FAIL on a
stale-or-absent preload plan** — it is 46 disk reads and a venue crackle away
from being a real preflight item, and it costs one `check-preload.sh` call,
exactly like the `parvagues-protect --check` fix.
- **Block size 1024 (~21 ms)** — high for livecoding. Worth tuning the PipeWire quantum.
**Cannot be answered until preload is warm** — see the entry above. Both
candidates measured on a cold cache rejected (1024: W/Q 1.300, +6,982 ERR in
60 s; 512: W/Q 1.390, +17,348), and W/Q > 1 at the *stock* setting means the
measurement was reading disk stalls, not DSP headroom. Re-run after a warm
boot before drawing any conclusion about quantum.
**Groundwork done 2026-09-07: `tools/latency-lens.py`.** Nothing had ever
chosen 1024 — there is no PipeWire config on this box at all
(`~/.config/pipewire/` and `/etc/pipewire/pipewire.conf.d/` both absent), so
it is the raw upstream default, which current guidance places in the
mixing/mastering bracket (no live input); live monitoring sits near 256
(5.3 ms period, ~10.7 ms round trip). It only became worth measuring once
scsynth reliably ran at SCHED_FIFO/90 — the guidance is blunt that without RT
scheduling every other tuning is marginal, and until `a828d5f` protect was
handing scsynth back at `SCHED_OTHER/0`.
The lens implements the documented method rather than a guess: accept the
LOWEST quantum holding **ERR delta 0** and **max W/Q < 0.75** under a
representative load, sampled from `pw-top -b`. It sets `clock.force-quantum`
at runtime and restores it; it never edits a config (`--print-config Q`
prints the drop-in to write when a choice is actually made).
**It refuses to run unless the default sink is muted, continuously.** A valid
quantum test needs the real hardware device driving the graph (a null sink is
driven by a software timer and answers far too optimistically), but it does
not need the speakers: mute sits downstream of all DSP. The mute is
re-checked ~1/s for the whole run and the sweep dies the instant it stops
holding — because a one-shot check is a snapshot, not a guarantee, and this
was proven twice in ten minutes (see the learnings below).
**Still to do:** the real acceptance run — `--dwell 600`, UMC202HD plugged
in, Ardour recording. USB has its own floor and the gig path is the UMC, so
no number measured on internal audio should be trusted for a venue.
- **Nothing on this rig counts SuperCollider `late` messages**, and they are the
single best timing-health signal a Tidal rig has. 136 of them on 2026-09-06
went unnoticed: `gig-log` records thermals, freq, cpu, gear, xrun deltas,
MIDI, track and eval — not `late`. Found only by grepping the journal by hand.
They came in **two bursts** (64 at 19:58, 72 at 20:03), values *decaying*
8.1 s → 0.65 s like a drained backlog, not the steady trickle that means
"latency too low". Both bursts sit against SC restarts, and the 20:03 one is
the watchdog bug fixed in `b956409`, caught in the act:
`20:03:35 scsynth GONE (3 polls) … restarting (1 prior in window)` six
seconds into a healthy boot. `late` belongs in gig-log's `s` record as a
delta, next to the xruns.
- **`s.latency = 0.3` is probably padding for a cause that was misdiagnosed.**
`start_and_midi.scd:349`, comment "increase this if you get late messages",
history `1` → `0.3` (`d008a31`) and never revisited. SC's documented default
is 0.2, and 0.3 s is what sits between ctrl+enter and hearing the change. The
documented reason to pad it — late messages — was two restart bursts, not
insufficient latency. **Not changed blind**: walking it down wants the `late`
counter above so the rollback signal is visible, plus PLN's ears. Cheap and
reversible once both exist.
- **No ld.so.conf.d preference for PipeWire's libjack**, so `libjack.so.0` resolves to
jackd2's. `pw-jack` works around it per-app; a system-wide preference would also make
qjackctl show the graph the rig is actually in, instead of an empty jackd2 one. Needs
sudo, and would leave jackd2 effectively unused — which is the point.
- **The headset mic auto-connects into `SuperCollider:in_1/2`** on every SC boot
(`SC_JACK_DEFAULT_INPUTS=system`). `check-audio-graph.sh` prunes it, but it recurs.
- **Re-run `check-audio-graph.sh` at the venue with the UMC plugged in** — it cannot
prove the last monitor hop until the interface exists (2 warnings tonight, both that).
## The jig, second half
- **Only the reference-clearing half of the fresh-jig plan is built.** The stem sweep
(copy -> verify -> clear, idempotent, at launch as well as close) is designed but not
implemented. Three open questions still need PLN: per-gig sessions vs one stable jig;
where swept stems land (`take-master.py`/`take-lens.py` should own it, not a new tool);
whether the local 53G ever comes back or the mirror stays the archive of record.
## Audio session — opened 2026-09-07 (evening, travel rig)
- **`perf-tray` should be drawn by the SHELL, not by Qt.** Domovoy's tray is
the cleaner system and the reason is architectural, not cosmetic: it uses
**pystray**, which on this session picks the AppIndicator/Ayatana backend
(its bus path is `/org/ayatana/NotificationItem/domovoy`), so GNOME draws the
menu from the exported DBusMenu — native look, native left-click, no palette
to maintain. Qt5's `QSystemTrayIcon` hardcodes StatusNotifierItem
`ItemIsMenu=false`, so the host will NEVER open our menu; the left-click
handler has to pop a Qt window at the cursor, and that window is ours to
style forever. System `python3` already has `gi` +
`AyatanaAppIndicator3-0.1` + `Notify-0.7` typelibs, so the port needs **no
new dependency**: AppIndicator3 for the icon (set_icon_full wants a file, so
the PIL sparkline gets written to a temp PNG per refresh — which is what
pystray does underneath anyway), Gtk.RadioMenuItem for the two axes,
GLib.timeout_add for the 2 s refresh, Notify for showMessage. ~200 of the
833 lines are the Qt shell; the gearbox logic is untouched. Do it on a quiet
day, not the night before a trip — 2026-09-07 shipped Fusion + a login-race
fix instead, which makes the current menu presentable rather than native.
- **The tools test suite has no interpreter that can run all of it.** Under the
pyenv `python3` (the shell default) three `test_lcxl3_display` tests fail
because `mido` is not installed there; `/usr/bin/python3`, which HAS mido and
is what every systemd unit runs, has no `pytest`. So the suite is green only
by looking away. Fix: `sudo apt install python3-pytest` (or a venv that uses
the system interpreter), then make the failure loud if the interpreter is the
wrong one — the pyenv-vs-system trap has now cost this rig three separate
wrong reports.
- **`fold-orbits` sums at unity, with no trim.** 14 orbits playing at once is
14x one orbit's level, so a full arrangement on headphones can clip where the
same set through Ardour's Master does not. Deliberate for now (no hidden gain
to explain later); if it bites, the honest fix is a `--trim` that says what it
did, not a silent -12 dB.
- **The fold does not follow the sink.** Plug headphones in and the default sink
changes; the links stay on the old node and the ears get silence. `--check`
catches it in one line and a re-run fixes it. A `--watch` mode that reconciles
on graph change would remove the step — but it is another daemon on a laptop
that already runs three reconcile loops, so measure the annoyance first.
## Audio session — opened 2026-09-07
### Afternoon — the undead server (closed 2026-09-07)
- ~~**Glitchy/laggy desktop audio (media keys, FIP)**~~ **ROOT-CAUSED + LAYERED
FIXES 2026-09-07 afternoon.** An scsynth whose pipewire-jack client loses its
startup race gets NO negotiated format (pw-top: quantum 0/rate 0,
self-driving, BUSY overflowing) and busy-spins its data loop at ~95% of a
core at SCHED_FIFO — while both the unit and the process look perfectly
healthy. Hours of that starved the SOF DSP's IPC until the firmware wedged
(`sof-audio ... IPC timeout` -110) and EVERY sink vanished. The overnight
instance burned 9h21m of CPU. Three occurrences in one day; same config also
booted clean at 1.5% — **a probabilistic race, so the durable protection is
detection + self-heal** (sc-watchdog spin detector, tests 12-15; rig-doctor
`scsynth orphan spin` + `SOF audio DSP` + `failed units` checks), plus one
real prevention: autoroute no longer prunes SC's links when Ardour is absent.
- **Wedged SOF DSP, no-root cure (verified live):** `timeout 5 speaker-test
-D hw:sofsoundwire,2 -c 2 -t sine` forces a runtime resume → driver
cold-boots the firmware (`IMR restore failed, trying to cold boot`) → then
`systemctl --user restart wireplumber` rebuilds sinks. No reboot needed.
- **mi global effects: measured LIGHT, acquitted.** 5 idle boots: off=1.7%,
all=1.3%, verb=1.7%, clouds=1.7%, ripples=2.0% of one core. Gated behind
`PV_MI_FX` (default all) so the next suspicion is a knob-flip, not a guess.
NOTE: no track in live/ or copycat/ uses their params yet — the ear test
(PLN, deferred) is still the only consumer.
- **FIP stuttering on Display-1 output:** two causes overlapped. The hard
stutter at 13:41-13:47 coincided exactly with a spinning scsynth (82% RT).
With SuperDirt stopped, mpv ran ERR 0 / µs waits over 20s — residual
hiccups are the FIP network feed / mpv buffering, not the graph.
- **Open:** the quantum sweep (needs SuperDirt DRIVEN + dense track; now
meaningful since the DSP-load red herring is resolved); watchdog spin
detector unproven against a REAL spin (only the fake) — next natural
occurrence is its live test, watch for `SPINNING` in the watchdog journal.
- 🔴 **`wireplumber` hitting `start-limit-hit` takes ALL audio out, and looks
like dead hardware.** Reported by PLN as "media keys dont work" + "tried
bluetooth headset couldnt get sound there". Diagnosis:
wireplumber failed (start-limit-hit)
Sinks: 100. Dummy Output <- the ONLY sink
0 bluez objects
pipewire is the graph; **wireplumber is the session manager that puts DEVICES
in the graph**. Without it there are no device nodes at all: every hardware
sink disappears, PipeWire invents `Dummy Output`, media keys act on nothing,
and a Bluetooth headset cannot appear however perfectly it pairs.
**The trap**: `StartLimitBurst=5` / `StartLimitIntervalUSec=5min` with
`Restart=on-failure`. Five restarts inside two minutes (10:25:01, 10:25:51,
10:26:07, 10:26:26, 10:27:05 — each `stopped by signal: Terminated` then
immediately `Started`, i.e. hand-issued `systemctl --user restart`) exhausted
it, and from then on **restarting again cannot work** — the journal says
`Start request repeated too quickly` and nothing starts. It recurred at
11:43-11:44 while this was being written, with
`systemctl[2200836]: Job for wireplumber.service failed` naming the client.
**Cure, and it is not another restart:**
systemctl --user reset-failed wireplumber && systemctl --user start wireplumber
**Third instance of this exact shape in one session** — `parvagues-sc` twice
(2026-09-06 23:17 and the boot-grace bug) and `wireplumber` here. On this box,
"X will not start" means check `reset-failed` BEFORE anything else.
- ~~**`rig-doctor` could not see a total audio outage**~~ **FIXED 2026-09-07.**
`check_pipewire` tested `pipewire` on PATH, `pw-jack` on PATH, and
`pipewire.service is-active` — and **`pipewire.service` was `active` for the
whole outage**, so the check was green while the box had no audio whatsoever.
Presence versus function, again. It now also asserts wireplumber is active AND
that the graph holds at least one real output (`alsa_output.*` /
`bluez_output.*`, never `auto_null`/Dummy), FAILs with the `reset-failed` cure
in the fix field, and WARNs if a dummy is present or the default is unset.
Caught a LIVE regression the first time it ran: `wireplumber is 'failed' — 0
real sink(s), 1 dummy`, while PLN was still retrying by hand.
Doctor now: **48 checks, 0 fail, 10 warn, 38 pass.**
- ~~**Bluetooth headset had no sound**~~ **FIXED 2026-09-07: it was muted.** Once
wireplumber was up, `bluez_card.88_C9_E8_9F_C1_49` was present with
`Active Profile: a2dp-sink` (High Fidelity Playback, SBC) — the correct
profile, never the problem. The sink read
`137. WH-1000XM5 [vol: 0.20 MUTED]`: per-device mute that wireplumber
persists in its own state, at 20% volume. `pactl set-sink-mute
bluez_output.88_C9_E8_9F_C1_49.1 0` → `Mute: no`. **Not caused by the
quantum/mute testing**: that ran at 00:10 when no BT device was connected, and
`clock.force-quantum` was verified back at 0.
Second half of "no sound there", left for PLN on purpose: the XM5 is **not the
default sink**, so nothing routes to it until selected —
`wpctl set-default <id>` or the sound tray. Switching a live default would
yank audio off the speakers mid-session.
NB the BT node id MOVES between calls while the device reconnects (105183 →
105373 within a minute), so address it by NAME, never by id.
- **`latency-lens --print-config` used to hand you the loaded gun.** Its "Apply"
line is `systemctl --user restart pipewire pipewire-pulse wireplumber`, which
run a few times in a row is exactly the trap above. It now says ONCE, explains
the start-limit consequence, gives the `reset-failed` cure, and points at the
runtime `pw-metadata` setting that needs no restart at all.
## Housekeeping
- **`rig-doctor` does not check LV2 plugins** — the same gap class as the PySide6 miss.
- ~~**`parvagues-protect` burns 10.7% of a core, forever, on a battery laptop.**~~
**FIXED 2026-09-06 (late), pending a root install.** Re-measured to 219.1 ms
per 2 s tick = **10.96% of a core with nothing running to protect**, so the
original figure was if anything generous. It was all discovery, none of it
protection: three `pgrep -x` at 59 ms each (pgrep reads cmdline for every
process on the box, and it was called once per target), plus `cat`, two
`chrt -p` reads and an unconditional `prlimit` per pid, each inside a command
substitution. Now answered entirely with bash builtins — one pass over
`/proc/*/comm`, sched read from `/proc/<pid>/stat` fields 40/41, `prlimit`
only when `/proc/<pid>/limits` says it is needed:
| | ms/tick | forks/tick | % of a core @2 s |
|---|---|---|---|
| before | 219.1 | 6 | 10.96% |
| after | 28.5 | **0** | **1.42%** |
The interval is untouched at 2 s deliberately — the 7.7× came from forks
alone, so the responsiveness that catches a restart before first sound was
not traded away. `PARVAGUES_PROTECT_INTERVAL=5` takes it to 0.57% if the
battery ever matters more than a 5 s window of non-RT audio after a restart.
What made it cheap was the *prefilter*: a `[sSaA]*` character class rejects
606 of 641 processes before bash's regex engine — by far the slowest thing in
reach — is consulted. Testing the regex on every comm costs 149 ms/tick,
**worse than the pgrep it replaced**. Verified differentially against the old
`pgrep` path with a decoy fleet (`ardour9`, `ArdourGUI`, `ardour-8.6` must
match; `ardour-decoy`, `sclang-notreally`, `scsynthx` and a `bash` whose path
contains "ardour" must not) and against `chrt -p`/`cat` on live scsynth+sclang.
**Still to do: `sudo tools/install-protect.sh`** — the repo has the fix, the
running daemon does not.
- ~~**`sc-watchdog` restarts SuperDirt on every clean start, and two starts in
five minutes leave the rig `failed` and unstartable.**~~ **FIXED 2026-09-06
(late).** `BOOT_GRACE_SECS=30` (injectable as `SCWD_BOOT_GRACE`): once the
miss threshold is reached, the loop asks how long the unit has actually been
active and holds off if that is under the grace. Misses keep counting during
it, so an expiring grace acts immediately rather than restarting a 6 s count.
Costs no extra fork — the per-poll `systemctl --user is-active` became one
`systemctl --user show -p ActiveState -p ActiveEnterTimestamp` — though not
nothing, measured rather than assumed: 7.0 ms → 9.4 ms, so +0.12% of a core.
The `date -d` parse runs only on the rare path where a restart was about to
happen.
**And while measuring that, the watchdog turned out to pay the same 59 ms
`pgrep -x` protect had just been cured of**, on every poll while SC is up —
3.0% of a core *while playing*. Replaced with `proc_alive()`, the same
stateless `/proc/*/comm` pass (52.9 ms → **20.8 ms**, zero forks,
differentially verified against `pgrep -x` on `scsynth`/`scsynthx`/
`myscsynth`/`sclang`/absent). Deliberately duplicated rather than shared:
parvagues-protect is copied to `/usr/local/bin` and runs as root, so it must
not source anything from a user-writable repo. Net: while playing 3.00% →
1.51%; while idle 0.35% → 0.47%, the `show` call being the price of the
correctness fix. Regression test added as case 6 in `tools/tests/test-sc-watchdog.sh`:
a fake unit that takes 10 s to produce its server, asserting no `GONE (` line
and no rate-limit spend. Suite: 11 passed, 0 failed.
Original diagnosis, kept because the shape recurs — found 2026-09-06 by
starting `parvagues-sc` for an unrelated test. `parvagues-sc.service` is
`Type=simple`, so it reports `active` the instant `sclang` execs — but
`scsynth` only appears ~8 s later, when SuperDirt boots the server. The
watchdog's main loop counts a miss every `POLL_SECS=2` while the unit is
active and scsynth is absent, and acts at `MISSES_TO_ACT=3`. **6 s < 8 s**, so
a perfectly healthy start always trips it. Journal, verbatim, from a fresh
start with nothing wrong: `23:17:12 scsynth GONE (3 polls) while
parvagues-sc.service is active — restarting (0 prior in window)` → `23:17:21
scsynth up after 8s` → `recovered`.
The consequence is the gig-night one: each start burns two of systemd's
`StartLimitBurst=3` (PLN's, then the watchdog's), so a *second* start inside
`StartLimitIntervalUSec=5min` hits the limit and the unit goes
`failed (result: start-limit-hit)` — SuperDirt then refuses to start at all
until `systemctl --user reset-failed parvagues-sc`. On stage that reads as
"the rig is dead and will not come back". Reproduced end to end tonight and
cleared with `reset-failed`; the box is back to `inactive/linked` as found.
Fix: give the watchdog a boot grace keyed to the unit's own
`ActiveEnterTimestamp` — do not count misses until the unit has been active
longer than SuperDirt's measured boot (~8-12 s; 25 s is a safe floor). It
already shells `systemctl --user is-active` every poll, so the timestamp is
one field on a call it is making anyway. `await_scsynth`'s
`BOOT_WAIT_SECS=100` covers the *post-restart* wait and was never wired to the
*initial* start, which is the whole bug.
- **The other two reconcile loops cost 5.3% of a core between them, and both can
be event-driven.** Measured 2026-09-06 on the same idle box:
`midi-autoconnect` **2.74%**, `tidal-ardour-autoroute` **2.53%**,
`parvagues-sc-watchdog` 0.48%, `parvagues-bridge` 0.07%, `midiviz` 0.09%.
Neither of the big two is slow per call (`aconnect -l` ~10 ms, `pw-link -l`
~11 ms); they just make 3-5 of them plus awk every 2 s to re-discover an
unchanged graph. Both have a real event source: `pw-link -m/--monitor` blocks
and prints on link/port change, and ALSA-seq exposes
`/proc/asound/seq/clients`, which a bash `read` can compare between ticks for
free. Reconcile on change instead of on a timer and each drops to ~0.1%.
Total idle rig cost today: **~14.5% of a core with nothing playing**; protect's
fix takes that to ~5.8%, and these two would take it under 1%.
**Probed 2026-09-06, so the next session starts from evidence, not hope:**
`pw-link -m -o` blocks, dumps current state prefixed `=`, and then emits on
change — that is the event source, confirmed by hand. (`pw-link -m -l`
printed nothing in 5 s; use `-o`/`-i`, not `-l`.) Shape:
`reconcile` once, then `pw-link -m -o | while read -r _; do <debounce>;
reconcile; done`, with `Restart=always` covering pw-link dying and a slow
fallback timer covering a missed event. **Deliberately NOT done tonight**:
both scripts are on the gig path and their own headers say a bug out here is
merely noisy only because the *boot* path was kept separate. Rewriting two
reconcile loops' control flow at midnight is how that stops being true.
This one wants its own session, with a `--dry-run` mode first, the way
sc-watchdog earned one.
- ~~**`rig-doctor`'s `parvagues-protect: PASS` tests for files, not for
function.**~~ **FIXED 2026-09-06 (late).** `check_protect()` now asks three
questions instead of one: are the files there, is the daemon *running* with
`cap_dac_override` in its live `CapEff` (read from `/proc/<MainPID>/status`,
no `capsh` dependency), and what does its own `--check` say about every
process that makes sound. Two rows now: `parvagues-protect` and
`parvagues-protect coverage`. Negative-tested against the live fault — pointed
at an unprivileged pid it reports FAIL and names the capability, and it
distinguishes "unverifiable" (pid gone → WARN) from "verified missing"
(`CapEff=0x0` → FAIL). One bug caught while writing it: the first version
used the `_unit_active()` helper, which is `--user`-scoped and appends
`.service` itself — it would have reported a healthy SYSTEM unit as inactive,
the same wrong-scope mistake as asking the wrong interpreter.
Original diagnosis: it tested for files, not function.
It checks that `bin` and `unit` exist — the exact pair that stayed true all
through 2026-09-06 while the daemon could not write a single `oom_score_adj`
(missing `CAP_DAC_OVERRIDE`, see the install entry above). Should shell out to
`/usr/local/bin/parvagues-protect --check`, which already reports per-process
truth and exits non-zero, and additionally assert the running daemon holds
`cap_dac_override`. Same gap class as the LV2 and PySide6 misses, but this one
guards the thing that keeps the music alive under memory pressure.
The missing `lsp-plugins-lv2` / `zynaddsubfx-lv2` only surfaced via Ardour's own dialog.
- **The rig now carries two Qt bindings** (PyQt5 for perf-tray, PySide6 for midiviz).
Documented in the doctor rather than unified; a PyQt5 port of midiviz is ~30 refs but
would move it off the Qt6 it was written and tested against.
- Untracked: `start_minimal.scd`, `sandbox/debug.tidal`. Also `BootTidal.hs.broken` and
`BootTidal.visuals.broken.hs` are still sitting in the repo root.
## Opened 2026-09-06 (second pass)
- ~~**The Ardour fader FEED is designed, not built.**~~ **BUILT 2026-09-06
night** (`3a91eba`). Two complementary sources: a forwarded CC counts as an
observation (the driver translated it), and Ardour echoes changes it did NOT
get from us. Own virtual port (`ParVagues LCXL3 FB`), JACK link discovered by
suffix, re-asserted every reconcile tick, re-links in <5 s across a restart.
**One leg is still unverified end to end**: no Ardour-side (GUI/mouse) fader
move was observed arriving, because nobody clicked one — the link is proven
up and the ingest path is proven correct, but the echo itself has not been
seen. Next time Ardour is open, drag a Tidal fader with the mouse and check
`journalctl --user -u lcxl3-driver -f` for an `ardour -> v2 #NN` line (needs
the driver run with `--verbose`, which the unit does not pass). If it never
arrives, the physical-move source still covers everything PLN asked for.
Related: the GLOBAL `"midi-feedback"` in `~/.config/ardour8/config` was
flipped 0->1 that night, but `<Protocol name="Generic MIDI" feedback="1">`
was already on, so that flip may have been unnecessary. Backup:
`config.pre-midifeedback-20260906-212052.bak`. Worth settling before assuming
either flag matters.
- ~~**`../tidal-ears/` is not cloned on this box.**~~ **DONE 2026-09-06** —
cloned from `git@git.nech.pl:pln/tidal-ears.git` (branch `main`, 1.4 M; the
3.2 G it occupies on xps22 is untracked analysis output, deliberately not
fetched). `tools/analyze_samples.py` resolves again.
**The gap itself is still open:** no install step fetches the sister repos
CLAUDE.md documents. `tools/rig-install.sh` should clone `tidal-ears` and
offer `visuals/scenes` + GLITCHWAVE, or `rig-doctor` should at minimum FAIL on
a dangling committed symlink — three separate manual fixes for one missing
line of setup.
- **pytest is absent from the rig interpreter.** The suite runs under pyenv
3.11.10 (which lacks `mido`), so 3 `test_lcxl3_display.py` tests fail for
environment reasons, while `/usr/bin/python3` — the interpreter every systemd
unit actually uses, and the one that HAS mido — cannot run pytest at all. The
rig's own hardware tests are therefore unrunnable where they matter. Same bug
class as the rig-doctor fix: asking the wrong interpreter.
- ~~**`pulsar-parvagues-hud` 78b1755 is committed but NOT pushed**~~ **PUSHED
2026-09-06** on explicit request (`368bee9..78b1755` → `main`). Carries the
scene-resolver fix without which no backdrop appears on this box.
# Rig sessions, 6–7 September 2026 — archived out of backlog.md
Lifted verbatim out of `backlog.md` on 2026-09-22 when PLN cleared its tail:
"deliberate killed tail was just a agent log or srth, capture that removed as
learning, nut then let me clean that backlog indeed was gboing tedious".
One correction worth keeping, because it changes how this file should be read:
it is **not** an agent log. It is two sessions of rig notes with real findings
in them, and six items that were still open when the tail was cut. Those are
listed first; the rest is the original text, unedited, in its original order.
## Still open when this was archived
* **`sudo tools/install-protect.sh` was never run.** The repo carries the fix
(`8f2c510`: the unit's `CapabilityBoundingSet` omitted `CAP_DAC_OVERRIDE`, so
every write to `/proc/<pid>/oom_score_adj` returned EACCES under an error
message that said "need root, have uid 0"). Until it is installed, the
OOM-protection daemon has never protected anything — scsynth and sclang read
`ok` only because they were already at the target, so no write was attempted.
Acceptance is a restart re-protecting on its own: `systemctl --user restart
parvagues-sc`, then `parvagues-protect --check`, expecting `oom:200->-1000`
naming the NEW pid.
* **The UMC202HD dwell run** (`--dwell 600`, interface plugged) was never done.
* **The Ardour jig's copy → verify → clear cycle** is designed, not implemented,
and three questions still need PLN — starting with per-gig sessions vs one
stable jig.
* **Tidal fader drag was never eyeballed** in an open Ardour.
* **No install step fetches the sister repos**, so a fresh box is not a rig.
* The tray item's `ItemIsMenu=false` means the host will never open our menu;
the left-click path is the only one that works.
## The original text
# Rig — opened 2026-09-06 (gig night on the new XPS24)
Floated during the session that got rose_rouge playing, not done. Full context:
`armada/tasks/completed-archive.md` (top entry) + `docs/2026-09-06-ardour-fresh-jig.md`.
## Visuals last mile — DONE 2026-09-06
Two corrections happened here, and both are worth more than the task was.
**First**, the plan in this slot ("one home `visuals/scenes/`, compress the
gifs, rewrite the refs repo-relative") proposed three things that were already
done or already decided against: the home exists and is the HUD's `scenesDir`
default (`package.json:101`); the gifs were compressed 2026-08-29
(`rose_bloom*.mp4`, 65-204 KB); and committing them was deliberately rejected
in `.gitignore:53-56` with the reason written in. A fence got kicked before
anyone read why it was there.
**Second**, the correction was itself half-wrong. "The refs work as designed,
nothing to rewrite" is true only on the box that authored the image.
`scenes.js:resolve()` tried an absolute spec and nothing else, and
`_preferDerived()` — the twin lookup — is only reached through a source that
EXISTS. So on the XPS24 every ref resolved to null with the mp4 sitting right
there, and the derived-twin mechanism was quietly conditional on being the
machine that made the gif.
Fixed in the HUD (`pulsar-parvagues-hud` 78b1755, **committed not pushed** —
that repo is on `main`): an absolute miss falls back to the stem against
scenesDir, the name `scene-ingest.cjs` already writes. SSOT design untouched;
it just stops requiring the SSOT to be local. Measured over the real corpus:
**0 of 5 distinct scene targets resolved before, 5 of 5 after**, and a target
with no twin still resolves to null.
Media synced: `visuals/scenes/`, 13 mp4s. Use `-H` — `decollage.mp4` and
`decolle.mp4` are hardlinks of one 120 MB file, so the 197 M source tree lands
as 124 M:
rsync -aH --partial xps22.local:'~/Work/Sound/Tidal/visuals/scenes/' visuals/scenes/
~~Still open, and a deliberate choice rather than an oversight~~ **SYNCED
2026-09-06.** `~/Work/Art/GLITCHWAVE` (3.1 G) is the SSOT *source* tree, needed
only to re-run `npm run scenes:ingest` for a NEW scene, never to play one. It
was deliberately left on xps22 on the grounds that authoring happens there —
PLN overruled that: "glitchwave cant hurt tbh". Fetched with `rsync -aH`, so
this box can now author scenes as well as play them, and the GIF `url:` lines
scattered through this file resolve locally.
rsync -aH xps22.local:'~/Work/Art/GLITCHWAVE/' ~/Work/Art/GLITCHWAVE/
## Sound / mix
PLN: "tidal 12 -inf is random i might turn on/off any fader its not signal
throughout perf it mioves always and i save random". Fader positions at save
time are arbitrary, so there was nothing to decide. Measured: **9 of 12**
faders had drifted, five to -inf. `gig-up.sh` now runs
`fader-baseline.py --restore` in the Ardour-closed window, full restore to the
2026-08-02 baseline. Nothing further unless the baseline itself should move
(`--capture`).
with its sudoers rule (`visudo -cf` parsed OK), so gig-up can set performance mode.
It had played at `gear: None, mode: silent` with the CPU capped at powersave.
immediately exposed a bug in the guard itself** — fixed in `8f2c510`. It ran as
root and still could not write `/proc/<pid>/oom_score_adj`: the unit's
`CapabilityBoundingSet` omitted `CAP_DAC_OVERRIDE`, which is what uid 0 uses to
bypass file permissions, and the target is mode 0644 owned by `pln`. Every write
returned EACCES under an error message that said "need root, have uid 0".
scsynth and sclang showed `ok` only because they were already at the target, so
no write was attempted — the daemon had never protected anything.
**Still to verify:** reinstall and watch a restart re-protect on its own, which
is the only test that proves it (`systemctl --user restart parvagues-sc`, then
`parvagues-protect --check` — expect `oom:200->-1000` naming the NEW pid).
off disk mid-set.** Found 2026-09-07 while trying to measure the quantum, and
it is the same failure `tools/check-preload.sh`'s header calls out as having
"debuted as a crackle at the venue" — except on xps22 the plan existed and
warmed the wrong list, while here there was no plan at all.
`start_and_midi.scd:259` tests `File.exists("preload.scd")` and silently
takes the other branch: *"preload: no preload.scd — lazy-loading samples on
demand."*
**Measured, not inferred.** Driving `rose_rouge` headlessly for ~3 min logged
**46** `reading soundfile as needed` lines — `rose:28.0`, `rose:29.0`,
`rose:30.0`, one every 5-12 s, each one a disk read landing on the audio
thread. In the same window the SuperCollider node's `pw-top` **W/Q reached
1.300 at the rig's stock 1024** and 1.390 at 512, i.e. exceeding its deadline
regardless of quantum. `check-preload.sh` then said it plainly: *"60 bank(s)
would be read from DISK on first play (crackle, mid-transition)"*.
**Fixed by running `tools/check-preload.sh --fix`** — 60 banks, all 16
setlist tracks, 7.5 KB, gitignored (correctly, it is generated). Verified at
the next boot: `preload: warming the set's samples…` /
`=== PRELOAD: 60/61 banks OK in 2.6 s ===`.
2026-09-07: `tools/gen_setlist.py`.** The preload now warms what PLN actually
PLAYS, computed from the canonical gig records rather than a hand-kept file,
so it cannot drift: **60 banks → 111**, covering **39 tracks** from the last
12 months instead of 16 from one August gig. 51 banks were one first-play
away from a disk read. `rose` — the bank behind all 41 measured lazy reads —
is covered.
Sources, in authority order: `<www>/content/lives/<year>/<slug>/tracks.json`
(canonical; carries both `date` and each track's exact repo-relative `file`)
for opal-festival-2026, montreuil-algorave, raise and bunker; then
`armada/tide-table/judge_specs/*_setlist_ear.json` for gigs whose tracks.json
is not built yet — that is where the **cosmicfest-2026** list lives, "THE
ground truth, 14 tracks", with `rose_rouge` at #5. Those carry no gig date, so
the tool includes them and SAYS SO rather than inventing one.
`check-preload.sh` now computes its own list instead of borrowing
`set-coherence.setlist_tracks()`, because the two ask different questions: a
cheat sheet wants TONIGHT's running order, a preload wants anything PLN might
play, and the cost asymmetry is total (a warmed unused bank costs boot
seconds; an unwarmed one costs a disk read on the audio thread at the venue).
Override with `PV_PRELOAD_SETLIST=path` for one specific gig, or
`PV_PRELOAD_MONTHS=6`.
**Measured cost of over-covering:** 124 banks eager, `122/124 OK in 5.0 s`
(was 60/61 in 2.6 s), scsynth RSS **4.8 G** — on a box with 62 G and 40 G
available, i.e. 8%. Fine. Two unresolved track names, both harmless:
`mafia` is an alias for `mafia_sans_serif.tidal`, already in via opal, and
`Outro Dub Siren` is a live improvisation with no file.
on PLN's explicit say-so** ("no reason to keep these imo"). **198 files, 0.3
MB**, every one confirmed by `file` magic before removal — the guard was
mechanical, not the name pattern, and 198/198 candidates came back
"AppleDouble encoded Macintosh file" with zero false positives.
69 were in playable banks: `rhadamanthe_fx` 227→189, `rhadamanthe_divers`
314→293, `rhadamanthe_vocal` 77→72, `rhadamanthe_melo` 32→30. The other 129
sat under `~/Work/Sound/Samples/baba/__MACOSX/`; the emptied dirs were
`rmdir`-ed (rmdir refuses a non-empty directory, so it cannot take data).
**Sample indices were NOT disturbed**: the `zz.` prefix sorted every one of
them after all real samples, which is presumably why someone renamed rather
than deleted. Only the counts moved. Boot is now free of the
`WARNING: File reading failed for path:` wall, and the preload reports
**124/124 banks OK in 5.2 s** where it used to report 122/124.
and the real bug was underneath it.** Three separate faults, in order of
discovery:
1. The banner's headline said "whitelist is STALE, regenerate it" when the
cause was unreadable files, and regenerating provably changed nothing. It
now says "N bank(s) DID NOT FULLY LOAD", separates the two causes
(`got == 0` = whole bank failed; `0 < got < expected` = folder changed OR
some files are not audio) and prints the `file`-based one-liner that tells
them apart. *(Correction to an earlier claim of mine: it always DID name
the offending banks — `bad.do { ... }` prints them. My journal grep was
filtering those lines out. Only the headline was wrong.)*
2. **`check-preload.sh` compared bank NAME SETS, never counts.** So a bank
whose contents changed was invisible: the plan said `rhadamanthe_vocal, 77`
against a folder holding 72 and the checker printed **ok**, while
SuperDirt's boot said `expected 77 files, got 72`. Worse, `--fix` refused
to regenerate, because by the name-set test nothing had changed. The
everyday way to hit this is not deleting files — it is **dropping a new
sample pack into an existing bank**, which PLN does, and every new file
would then lazy-load mid-set under a green check. Now compares
`name count` pairs and reports `plan N -> disk M` per drifted bank.
3. **The bank-name regex was case-blind** — `[a-z0-9_]+` truncated `rampleA0`
to `rample` and collapsed every `vocalOoh1`/`vocalScatJ`/... into one
`vocal`. That is why the checker said **111 banks** while SuperDirt loaded
**124**: 124 was always the truth, and the shortfall was the regex, which
was wrong in both directions — undercounting the plan AND inventing two
bank names that do not exist. Found by asking why two numbers never
matched instead of assuming one rounded the other. Fixed to
`[A-Za-z0-9_]+`; and because `bank_counts()` inherited it, drift in any
camelCase bank had still been invisible after fault 2 was fixed.
**Now trustable, three ways agreeing**: `check-preload.sh` says 124, the boot
says `124/124 banks OK`, an independent Python parse says 124 names / 124
pairs / no gap. Drift detection tested in both directions on a camelCase bank
(`vocalOoh1` plan 9 vs disk 13 → STALE exit 1; restored → ok exit 0, plan
byte-identical after the test).
three banks PLN plays.** `rhadamanthe_fx` (38), `rhadamanthe_divers` (21),
`rhadamanthe_vocal` (5) — files named `zz._<original>.wav`, 4096 bytes each,
and `file` calls them exactly what they are: "AppleDouble encoded Macintosh
file". They are macOS metadata, not audio. SuperCollider fails to read them
(that is the `122/124` above and a wall of `WARNING: File reading failed for
path:` at every boot), and because they are counted as files in the bank
**they occupy sample indices** — so `s "rhadamanthe_fx:n"` is silent for 38
values of n, and `bank_file_count` overcounts. Another 129 sit under
`~/Work/Sound/Samples/baba/__MACOSX/`, which is extraction junk rather than a
playable bank.
**NOT touched**: the `zz.` prefix looks like a deliberate rename to sort them
last rather than delete them, which is a fence with a reason until PLN says
otherwise. Deleting them is his call; `find <bank> -name 'zz._*' -delete`
after a look is the whole fix, and it would also silence the boot warnings
and un-shift the indices.
(unchanged): `122/124 banks OK` then `=== PRELOAD COUNT MISMATCH — whitelist
is STALE, regenerate it ===` on a plan generated sixty seconds earlier. Now
we know WHY the counts differ — the 2 misses are the AppleDouble banks above,
which the generator counts and the server cannot load — so the banner is
reporting a real fact with the wrong diagnosis. It should say "N banks failed
to LOAD" and name them, not "regenerate the whitelist", which does nothing.
the wrong set.** Re-running the
same load after the fix still logged 41 lazy reads — and **all 41 were the
`rose` bank**, because `rose_rouge` is NOT in `armada/setlist_opal2026.txt`.
Zero of the 60 preloaded banks lazy-loaded, and SC logged ZERO `late`
messages in that window, so the mechanism is provably
working; it is warming the 16-track OPAL set while the track this box has
actually been playing (see the rig-state memory) is outside it.
This is `check-preload.sh`'s own documented failure one layer up: there, the
plan was stale because two tools disagreed about the setlist (10 vs 13
tracks). Here the plan matches its setlist exactly and **the setlist is stale
relative to what gets played**. A freshness check that compares the plan to
the setlist cannot see this — it needs to compare against what is actually
loaded. Cheapest honest fix: add current material to the setlist (or a
second "current" list the generator also reads), then regenerate.
STALE, regenerate it`** on a plan generated sixty seconds earlier. So either
the generator counts one bank the server cannot load, or the two count
different things. Worth ten minutes: the mismatch banner is a real signal that
currently cries wolf on a fresh plan, and a warning that fires when nothing is
wrong is a warning nobody reads on the night it matters.
with one `s.sync`.** `global_mi_verb2` / `global_mi_clouds2` /
`global_mi_ripples2` were reported `SynthDef not found` on every clean boot —
**84 error lines**, i.e. 3 defs × 14 orbits × 2 lines. So **every orbit's
MiVerb, MiClouds and MiRipples send was a dead node for the whole session**:
`# verbwet`, `# cloudswet`, `# ripplesreson` did nothing, silently, because a
missing global effect is an absence of effect rather than a noise you notice.
**mi-UGens was installed perfectly all along** — 11 `.so` + 11 classes at
`~/.local/share/SuperCollider/Extensions/mi-UGens` — and the SynthDefs were
defined correctly at `start_and_midi.scd:273-296`. The bug was ordering:
`SynthDef(...).add` is ASYNCHRONOUS (compiles locally, sends `/d_recv`,
returns) and the very next statement, `~dirt.orbits.do { ... x.initNodeTree }`,
sends `/s_new` for those names. The `/s_new` overtook the `/d_recv`. One
`s.sync` before the registration, legal because we are inside
`s.waitForBoot`'s Routine — the idiom was already sitting commented out at
line 236. Verified: 84 errors before, **0 after**, same file one line apart.
(`spectral-freeze` above does NOT need it: it replaces a def whose synths are
built per event, long after `/d_recv` lands.)
where there was one presence test:
* `mi-UGens extension` enumerates the 11 plugins **and** the 11 classes we
actually use, and WARNs naming whichever is absent — `d.is_dir()` is true
whether the directory holds a working install or a stale README.
* `SuperDirt boot errors` (new, under SERVICES) scopes the journal to the
CURRENT boot and FAILs on `SynthDef ... not found` / `FAILURE IN SERVER`,
naming the distinct SynthDefs rather than repeating one line per orbit.
The question it asks is generic, so it catches the next async-ordering bug
too, not just this one. WARNs (not FAILs) when parvagues-sc is down, since
on-demand is the correct state before a set.
Both tested in BOTH directions: the boot-error check FAILs on the recorded
pre-fix journal window (naming all three defs) and PASSes on the post-fix
boot. That test is also what caught `re` never having been imported in
`rig-doctor.py` — my new code was its first user, so the check would have
crashed the doctor the first time it found something.
Doctor now: **47 checks, 0 fail, 10 warn, 37 pass.**
**The mechanism was never broken**: `gig-up.sh:182` already calls the checker
and auto-fixes on a non-zero exit, and the checker does exit 1 correctly.
`gig-up.sh` simply has not been run on this box since the setlist grew. So
the gap is not the code, it is that the gig-up path is the only thing that
regenerates the plan and nothing else notices. **`rig-doctor` should FAIL on a
stale-or-absent preload plan** — it is 46 disk reads and a venue crackle away
from being a real preflight item, and it costs one `check-preload.sh` call,
exactly like the `parvagues-protect --check` fix.
**Cannot be answered until preload is warm** — see the entry above. Both
candidates measured on a cold cache rejected (1024: W/Q 1.300, +6,982 ERR in
60 s; 512: W/Q 1.390, +17,348), and W/Q > 1 at the *stock* setting means the
measurement was reading disk stalls, not DSP headroom. Re-run after a warm
boot before drawing any conclusion about quantum.
**Groundwork done 2026-09-07: `tools/latency-lens.py`.** Nothing had ever
chosen 1024 — there is no PipeWire config on this box at all
(`~/.config/pipewire/` and `/etc/pipewire/pipewire.conf.d/` both absent), so
it is the raw upstream default, which current guidance places in the
mixing/mastering bracket (no live input); live monitoring sits near 256
(5.3 ms period, ~10.7 ms round trip). It only became worth measuring once
scsynth reliably ran at SCHED_FIFO/90 — the guidance is blunt that without RT
scheduling every other tuning is marginal, and until `a828d5f` protect was
handing scsynth back at `SCHED_OTHER/0`.
The lens implements the documented method rather than a guess: accept the
LOWEST quantum holding **ERR delta 0** and **max W/Q < 0.75** under a
representative load, sampled from `pw-top -b`. It sets `clock.force-quantum`
at runtime and restores it; it never edits a config (`--print-config Q`
prints the drop-in to write when a choice is actually made).
**It refuses to run unless the default sink is muted, continuously.** A valid
quantum test needs the real hardware device driving the graph (a null sink is
driven by a software timer and answers far too optimistically), but it does
not need the speakers: mute sits downstream of all DSP. The mute is
re-checked ~1/s for the whole run and the sweep dies the instant it stops
holding — because a one-shot check is a snapshot, not a guarantee, and this
was proven twice in ten minutes (see the learnings below).
**Still to do:** the real acceptance run — `--dwell 600`, UMC202HD plugged
in, Ardour recording. USB has its own floor and the gig path is the UMC, so
no number measured on internal audio should be trusted for a venue.
single best timing-health signal a Tidal rig has. 136 of them on 2026-09-06
went unnoticed: `gig-log` records thermals, freq, cpu, gear, xrun deltas,
MIDI, track and eval — not `late`. Found only by grepping the journal by hand.
They came in **two bursts** (64 at 19:58, 72 at 20:03), values *decaying*
8.1 s → 0.65 s like a drained backlog, not the steady trickle that means
"latency too low". Both bursts sit against SC restarts, and the 20:03 one is
the watchdog bug fixed in `b956409`, caught in the act:
`20:03:35 scsynth GONE (3 polls) … restarting (1 prior in window)` six
seconds into a healthy boot. `late` belongs in gig-log's `s` record as a
delta, next to the xruns.
`start_and_midi.scd:349`, comment "increase this if you get late messages",
history `1` → `0.3` (`d008a31`) and never revisited. SC's documented default
is 0.2, and 0.3 s is what sits between ctrl+enter and hearing the change. The
documented reason to pad it — late messages — was two restart bursts, not
insufficient latency. **Not changed blind**: walking it down wants the `late`
counter above so the rollback signal is visible, plus PLN's ears. Cheap and
reversible once both exist.
jackd2's. `pw-jack` works around it per-app; a system-wide preference would also make
qjackctl show the graph the rig is actually in, instead of an empty jackd2 one. Needs
sudo, and would leave jackd2 effectively unused — which is the point.
(`SC_JACK_DEFAULT_INPUTS=system`). `check-audio-graph.sh` prunes it, but it recurs.
prove the last monitor hop until the interface exists (2 warnings tonight, both that).
## The jig, second half
(copy -> verify -> clear, idempotent, at launch as well as close) is designed but not
implemented. Three open questions still need PLN: per-gig sessions vs one stable jig;
where swept stems land (`take-master.py`/`take-lens.py` should own it, not a new tool);
whether the local 53G ever comes back or the mirror stays the archive of record.
## Audio session — opened 2026-09-07 (evening, travel rig)
the cleaner system and the reason is architectural, not cosmetic: it uses
**pystray**, which on this session picks the AppIndicator/Ayatana backend
(its bus path is `/org/ayatana/NotificationItem/domovoy`), so GNOME draws the
menu from the exported DBusMenu — native look, native left-click, no palette
to maintain. Qt5's `QSystemTrayIcon` hardcodes StatusNotifierItem
`ItemIsMenu=false`, so the host will NEVER open our menu; the left-click
handler has to pop a Qt window at the cursor, and that window is ours to
style forever. System `python3` already has `gi` +
`AyatanaAppIndicator3-0.1` + `Notify-0.7` typelibs, so the port needs **no
new dependency**: AppIndicator3 for the icon (set_icon_full wants a file, so
the PIL sparkline gets written to a temp PNG per refresh — which is what
pystray does underneath anyway), Gtk.RadioMenuItem for the two axes,
GLib.timeout_add for the 2 s refresh, Notify for showMessage. ~200 of the
833 lines are the Qt shell; the gearbox logic is untouched. Do it on a quiet
day, not the night before a trip — 2026-09-07 shipped Fusion + a login-race
fix instead, which makes the current menu presentable rather than native.
pyenv `python3` (the shell default) three `test_lcxl3_display` tests fail
because `mido` is not installed there; `/usr/bin/python3`, which HAS mido and
is what every systemd unit runs, has no `pytest`. So the suite is green only
by looking away. Fix: `sudo apt install python3-pytest` (or a venv that uses
the system interpreter), then make the failure loud if the interpreter is the
wrong one — the pyenv-vs-system trap has now cost this rig three separate
wrong reports.
14x one orbit's level, so a full arrangement on headphones can clip where the
same set through Ardour's Master does not. Deliberate for now (no hidden gain
to explain later); if it bites, the honest fix is a `--trim` that says what it
did, not a silent -12 dB.
changes; the links stay on the old node and the ears get silence. `--check`
catches it in one line and a re-run fixes it. A `--watch` mode that reconciles
on graph change would remove the step — but it is another daemon on a laptop
that already runs three reconcile loops, so measure the annoyance first.
## Audio session — opened 2026-09-07
### Afternoon — the undead server (closed 2026-09-07)
FIXES 2026-09-07 afternoon.** An scsynth whose pipewire-jack client loses its
startup race gets NO negotiated format (pw-top: quantum 0/rate 0,
self-driving, BUSY overflowing) and busy-spins its data loop at ~95% of a
core at SCHED_FIFO — while both the unit and the process look perfectly
healthy. Hours of that starved the SOF DSP's IPC until the firmware wedged
(`sof-audio ... IPC timeout` -110) and EVERY sink vanished. The overnight
instance burned 9h21m of CPU. Three occurrences in one day; same config also
booted clean at 1.5% — **a probabilistic race, so the durable protection is
detection + self-heal** (sc-watchdog spin detector, tests 12-15; rig-doctor
`scsynth orphan spin` + `SOF audio DSP` + `failed units` checks), plus one
real prevention: autoroute no longer prunes SC's links when Ardour is absent.
-D hw:sofsoundwire,2 -c 2 -t sine` forces a runtime resume → driver
cold-boots the firmware (`IMR restore failed, trying to cold boot`) → then
`systemctl --user restart wireplumber` rebuilds sinks. No reboot needed.
all=1.3%, verb=1.7%, clouds=1.7%, ripples=2.0% of one core. Gated behind
`PV_MI_FX` (default all) so the next suspicion is a knob-flip, not a guess.
NOTE: no track in live/ or copycat/ uses their params yet — the ear test
(PLN, deferred) is still the only consumer.
stutter at 13:41-13:47 coincided exactly with a spinning scsynth (82% RT).
With SuperDirt stopped, mpv ran ERR 0 / µs waits over 20s — residual
hiccups are the FIP network feed / mpv buffering, not the graph.
meaningful since the DSP-load red herring is resolved); watchdog spin
detector unproven against a REAL spin (only the fake) — next natural
occurrence is its live test, watch for `SPINNING` in the watchdog journal.
like dead hardware.** Reported by PLN as "media keys dont work" + "tried
bluetooth headset couldnt get sound there". Diagnosis:
wireplumber failed (start-limit-hit)
Sinks: 100. Dummy Output <- the ONLY sink
0 bluez objects
pipewire is the graph; **wireplumber is the session manager that puts DEVICES
in the graph**. Without it there are no device nodes at all: every hardware
sink disappears, PipeWire invents `Dummy Output`, media keys act on nothing,
and a Bluetooth headset cannot appear however perfectly it pairs.
**The trap**: `StartLimitBurst=5` / `StartLimitIntervalUSec=5min` with
`Restart=on-failure`. Five restarts inside two minutes (10:25:01, 10:25:51,
10:26:07, 10:26:26, 10:27:05 — each `stopped by signal: Terminated` then
immediately `Started`, i.e. hand-issued `systemctl --user restart`) exhausted
it, and from then on **restarting again cannot work** — the journal says
`Start request repeated too quickly` and nothing starts. It recurred at
11:43-11:44 while this was being written, with
`systemctl[2200836]: Job for wireplumber.service failed` naming the client.
**Cure, and it is not another restart:**
systemctl --user reset-failed wireplumber && systemctl --user start wireplumber
**Third instance of this exact shape in one session** — `parvagues-sc` twice
(2026-09-06 23:17 and the boot-grace bug) and `wireplumber` here. On this box,
"X will not start" means check `reset-failed` BEFORE anything else.
`check_pipewire` tested `pipewire` on PATH, `pw-jack` on PATH, and
`pipewire.service is-active` — and **`pipewire.service` was `active` for the
whole outage**, so the check was green while the box had no audio whatsoever.
Presence versus function, again. It now also asserts wireplumber is active AND
that the graph holds at least one real output (`alsa_output.*` /
`bluez_output.*`, never `auto_null`/Dummy), FAILs with the `reset-failed` cure
in the fix field, and WARNs if a dummy is present or the default is unset.
Caught a LIVE regression the first time it ran: `wireplumber is 'failed' — 0
real sink(s), 1 dummy`, while PLN was still retrying by hand.
Doctor now: **48 checks, 0 fail, 10 warn, 38 pass.**
wireplumber was up, `bluez_card.88_C9_E8_9F_C1_49` was present with
`Active Profile: a2dp-sink` (High Fidelity Playback, SBC) — the correct
profile, never the problem. The sink read
`137. WH-1000XM5 [vol: 0.20 MUTED]`: per-device mute that wireplumber
persists in its own state, at 20% volume. `pactl set-sink-mute
bluez_output.88_C9_E8_9F_C1_49.1 0` → `Mute: no`. **Not caused by the
quantum/mute testing**: that ran at 00:10 when no BT device was connected, and
`clock.force-quantum` was verified back at 0.
Second half of "no sound there", left for PLN on purpose: the XM5 is **not the
default sink**, so nothing routes to it until selected —
`wpctl set-default <id>` or the sound tray. Switching a live default would
yank audio off the speakers mid-session.
NB the BT node id MOVES between calls while the device reconnects (105183 →
105373 within a minute), so address it by NAME, never by id.
line is `systemctl --user restart pipewire pipewire-pulse wireplumber`, which
run a few times in a row is exactly the trap above. It now says ONCE, explains
the start-limit consequence, gives the `reset-failed` cure, and points at the
runtime `pw-metadata` setting that needs no restart at all.
## Housekeeping
**FIXED 2026-09-06 (late), pending a root install.** Re-measured to 219.1 ms
per 2 s tick = **10.96% of a core with nothing running to protect**, so the
original figure was if anything generous. It was all discovery, none of it
protection: three `pgrep -x` at 59 ms each (pgrep reads cmdline for every
process on the box, and it was called once per target), plus `cat`, two
`chrt -p` reads and an unconditional `prlimit` per pid, each inside a command
substitution. Now answered entirely with bash builtins — one pass over
`/proc/*/comm`, sched read from `/proc/<pid>/stat` fields 40/41, `prlimit`
only when `/proc/<pid>/limits` says it is needed:
| | ms/tick | forks/tick | % of a core @2 s |
|---|---|---|---|
| before | 219.1 | 6 | 10.96% |
| after | 28.5 | **0** | **1.42%** |
The interval is untouched at 2 s deliberately — the 7.7× came from forks
alone, so the responsiveness that catches a restart before first sound was
not traded away. `PARVAGUES_PROTECT_INTERVAL=5` takes it to 0.57% if the
battery ever matters more than a 5 s window of non-RT audio after a restart.
What made it cheap was the *prefilter*: a `[sSaA]*` character class rejects
606 of 641 processes before bash's regex engine — by far the slowest thing in
reach — is consulted. Testing the regex on every comm costs 149 ms/tick,
**worse than the pgrep it replaced**. Verified differentially against the old
`pgrep` path with a decoy fleet (`ardour9`, `ArdourGUI`, `ardour-8.6` must
match; `ardour-decoy`, `sclang-notreally`, `scsynthx` and a `bash` whose path
contains "ardour" must not) and against `chrt -p`/`cat` on live scsynth+sclang.
**Still to do: `sudo tools/install-protect.sh`** — the repo has the fix, the
running daemon does not.
five minutes leave the rig `failed` and unstartable.**~~ **FIXED 2026-09-06
(late).** `BOOT_GRACE_SECS=30` (injectable as `SCWD_BOOT_GRACE`): once the
miss threshold is reached, the loop asks how long the unit has actually been
active and holds off if that is under the grace. Misses keep counting during
it, so an expiring grace acts immediately rather than restarting a 6 s count.
Costs no extra fork — the per-poll `systemctl --user is-active` became one
`systemctl --user show -p ActiveState -p ActiveEnterTimestamp` — though not
nothing, measured rather than assumed: 7.0 ms → 9.4 ms, so +0.12% of a core.
The `date -d` parse runs only on the rare path where a restart was about to
happen.
**And while measuring that, the watchdog turned out to pay the same 59 ms
`pgrep -x` protect had just been cured of**, on every poll while SC is up —
3.0% of a core *while playing*. Replaced with `proc_alive()`, the same
stateless `/proc/*/comm` pass (52.9 ms → **20.8 ms**, zero forks,
differentially verified against `pgrep -x` on `scsynth`/`scsynthx`/
`myscsynth`/`sclang`/absent). Deliberately duplicated rather than shared:
parvagues-protect is copied to `/usr/local/bin` and runs as root, so it must
not source anything from a user-writable repo. Net: while playing 3.00% →
1.51%; while idle 0.35% → 0.47%, the `show` call being the price of the
correctness fix. Regression test added as case 6 in `tools/tests/test-sc-watchdog.sh`:
a fake unit that takes 10 s to produce its server, asserting no `GONE (` line
and no rate-limit spend. Suite: 11 passed, 0 failed.
Original diagnosis, kept because the shape recurs — found 2026-09-06 by
starting `parvagues-sc` for an unrelated test. `parvagues-sc.service` is
`Type=simple`, so it reports `active` the instant `sclang` execs — but
`scsynth` only appears ~8 s later, when SuperDirt boots the server. The
watchdog's main loop counts a miss every `POLL_SECS=2` while the unit is
active and scsynth is absent, and acts at `MISSES_TO_ACT=3`. **6 s < 8 s**, so
a perfectly healthy start always trips it. Journal, verbatim, from a fresh
start with nothing wrong: `23:17:12 scsynth GONE (3 polls) while
parvagues-sc.service is active — restarting (0 prior in window)` → `23:17:21
scsynth up after 8s` → `recovered`.
The consequence is the gig-night one: each start burns two of systemd's
`StartLimitBurst=3` (PLN's, then the watchdog's), so a *second* start inside
`StartLimitIntervalUSec=5min` hits the limit and the unit goes
`failed (result: start-limit-hit)` — SuperDirt then refuses to start at all
until `systemctl --user reset-failed parvagues-sc`. On stage that reads as
"the rig is dead and will not come back". Reproduced end to end tonight and
cleared with `reset-failed`; the box is back to `inactive/linked` as found.
Fix: give the watchdog a boot grace keyed to the unit's own
`ActiveEnterTimestamp` — do not count misses until the unit has been active
longer than SuperDirt's measured boot (~8-12 s; 25 s is a safe floor). It
already shells `systemctl --user is-active` every poll, so the timestamp is
one field on a call it is making anyway. `await_scsynth`'s
`BOOT_WAIT_SECS=100` covers the *post-restart* wait and was never wired to the
*initial* start, which is the whole bug.
be event-driven.** Measured 2026-09-06 on the same idle box:
`midi-autoconnect` **2.74%**, `tidal-ardour-autoroute` **2.53%**,
`parvagues-sc-watchdog` 0.48%, `parvagues-bridge` 0.07%, `midiviz` 0.09%.
Neither of the big two is slow per call (`aconnect -l` ~10 ms, `pw-link -l`
~11 ms); they just make 3-5 of them plus awk every 2 s to re-discover an
unchanged graph. Both have a real event source: `pw-link -m/--monitor` blocks
and prints on link/port change, and ALSA-seq exposes
`/proc/asound/seq/clients`, which a bash `read` can compare between ticks for
free. Reconcile on change instead of on a timer and each drops to ~0.1%.
Total idle rig cost today: **~14.5% of a core with nothing playing**; protect's
fix takes that to ~5.8%, and these two would take it under 1%.
**Probed 2026-09-06, so the next session starts from evidence, not hope:**
`pw-link -m -o` blocks, dumps current state prefixed `=`, and then emits on
change — that is the event source, confirmed by hand. (`pw-link -m -l`
printed nothing in 5 s; use `-o`/`-i`, not `-l`.) Shape:
`reconcile` once, then `pw-link -m -o | while read -r _; do <debounce>;
reconcile; done`, with `Restart=always` covering pw-link dying and a slow
fallback timer covering a missed event. **Deliberately NOT done tonight**:
both scripts are on the gig path and their own headers say a bug out here is
merely noisy only because the *boot* path was kept separate. Rewriting two
reconcile loops' control flow at midnight is how that stops being true.
This one wants its own session, with a `--dry-run` mode first, the way
sc-watchdog earned one.
function.**~~ **FIXED 2026-09-06 (late).** `check_protect()` now asks three
questions instead of one: are the files there, is the daemon *running* with
`cap_dac_override` in its live `CapEff` (read from `/proc/<MainPID>/status`,
no `capsh` dependency), and what does its own `--check` say about every
process that makes sound. Two rows now: `parvagues-protect` and
`parvagues-protect coverage`. Negative-tested against the live fault — pointed
at an unprivileged pid it reports FAIL and names the capability, and it
distinguishes "unverifiable" (pid gone → WARN) from "verified missing"
(`CapEff=0x0` → FAIL). One bug caught while writing it: the first version
used the `_unit_active()` helper, which is `--user`-scoped and appends
`.service` itself — it would have reported a healthy SYSTEM unit as inactive,
the same wrong-scope mistake as asking the wrong interpreter.
Original diagnosis: it tested for files, not function.
It checks that `bin` and `unit` exist — the exact pair that stayed true all
through 2026-09-06 while the daemon could not write a single `oom_score_adj`
(missing `CAP_DAC_OVERRIDE`, see the install entry above). Should shell out to
`/usr/local/bin/parvagues-protect --check`, which already reports per-process
truth and exits non-zero, and additionally assert the running daemon holds
`cap_dac_override`. Same gap class as the LV2 and PySide6 misses, but this one
guards the thing that keeps the music alive under memory pressure.
The missing `lsp-plugins-lv2` / `zynaddsubfx-lv2` only surfaced via Ardour's own dialog.
Documented in the doctor rather than unified; a PyQt5 port of midiviz is ~30 refs but
would move it off the Qt6 it was written and tested against.
`BootTidal.visuals.broken.hs` are still sitting in the repo root.
## Opened 2026-09-06 (second pass)
night** (`3a91eba`). Two complementary sources: a forwarded CC counts as an
observation (the driver translated it), and Ardour echoes changes it did NOT
get from us. Own virtual port (`ParVagues LCXL3 FB`), JACK link discovered by
suffix, re-asserted every reconcile tick, re-links in <5 s across a restart.
**One leg is still unverified end to end**: no Ardour-side (GUI/mouse) fader
move was observed arriving, because nobody clicked one — the link is proven
up and the ingest path is proven correct, but the echo itself has not been
seen. Next time Ardour is open, drag a Tidal fader with the mouse and check
`journalctl --user -u lcxl3-driver -f` for an `ardour -> v2 #NN` line (needs
the driver run with `--verbose`, which the unit does not pass). If it never
arrives, the physical-move source still covers everything PLN asked for.
Related: the GLOBAL `"midi-feedback"` in `~/.config/ardour8/config` was
flipped 0->1 that night, but `<Protocol name="Generic MIDI" feedback="1">`
was already on, so that flip may have been unnecessary. Backup:
`config.pre-midifeedback-20260906-212052.bak`. Worth settling before assuming
either flag matters.
cloned from `git@git.nech.pl:pln/tidal-ears.git` (branch `main`, 1.4 M; the
3.2 G it occupies on xps22 is untracked analysis output, deliberately not
fetched). `tools/analyze_samples.py` resolves again.
**The gap itself is still open:** no install step fetches the sister repos
CLAUDE.md documents. `tools/rig-install.sh` should clone `tidal-ears` and
offer `visuals/scenes` + GLITCHWAVE, or `rig-doctor` should at minimum FAIL on
a dangling committed symlink — three separate manual fixes for one missing
line of setup.
3.11.10 (which lacks `mido`), so 3 `test_lcxl3_display.py` tests fail for
environment reasons, while `/usr/bin/python3` — the interpreter every systemd
unit actually uses, and the one that HAS mido — cannot run pytest at all. The
rig's own hardware tests are therefore unrunnable where they matter. Same bug
class as the rig-doctor fix: asking the wrong interpreter.
2026-09-06** on explicit request (`368bee9..78b1755` → `main`). Carries the
scene-resolver fix without which no backdrop appears on this box.
Markdown is supported
0% or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment