Commit 7360021e by PLN (Algolia)

gig: two writers re-arm the Ardour FIFO, and the AC has to test both

perf-audio:565 does the same chrt -f -p 80 on Ardour that parvagues-protect
does, plus 573 for every child at 79, and perf-watch reasserts it on any
drift by itself -- the Bridge journal caught it doing exactly that at
09:20:15 on Sep 23. So an AC that only tests the pv-protect edit passes at
home and fails at the venue.

Also marks the epistemic line, because the next reader will act on this
hours before a gig: the xrun counts and their attribution are measured, the
kill mechanism is consistent with every measurement but was never
reproduced, and it cannot be reproduced on demand -- it needs one >200ms
uninterrupted GUI burst. What makes the xrun half plausible is that
sched_rt_runtime_us sits at the default 950000/1000000, so RT tasks passing
95% of a period get every RT task on that runqueue throttled for the
remaining 50ms, which is 2.3 quanta at 1024/48k.
parent 2726d1dc
...@@ -604,8 +604,17 @@ costs three things: ...@@ -604,8 +604,17 @@ costs three things:
### Thu morning — in this order ### Thu morning — in this order
- [ ] **1. Make Ardour survive its own launch.** Blocks every Ardour-side item - [ ] **1. Make Ardour survive its own launch.** Blocks every Ardour-side item
below. Two candidate fixes, both one line, PLN's call (the file is below. **TWO root files re-arm this, not one** — fixing either alone leaves
root-owned — `sudoedit /usr/local/bin/parvagues-protect`, then the bug live:
- `/usr/local/bin/parvagues-protect`, the `TARGETS` line (`ardour:80`),
re-applied every 2 s;
- `/usr/local/sbin/perf-audio:565` — `chrt -f -p 80 $ARDOUR`, plus
**line 573 `chrt -f -p 79 $child` for every child**. `perf-watch`
reasserts this on its own whenever it sees drift (seen in the Bridge
journal at 09:20:15 on Sep 23: "drift: reasserted standard" →
"✓ Set Ardour (PID …) to real-time priority 80"). So an AC that only
tests pv-protect will pass at home and fail at the venue.
Two candidate fixes, both one line each, PLN's call (`sudoedit` both, then
`sudo systemctl restart parvagues-protect`): `sudo systemctl restart parvagues-protect`):
- **(a) Drop the FIFO, keep the OOM shield.** `ardour:80` → protect - **(a) Drop the FIFO, keep the OOM shield.** `ardour:80` → protect
`oom_score_adj=-1000` only. Justification: `AudioEngine 1` already runs `oom_score_adj=-1000` only. Justification: `AudioEngine 1` already runs
...@@ -618,8 +627,22 @@ costs three things: ...@@ -618,8 +627,22 @@ costs three things:
AC for either: launch from the **desktop icon** (the path that failed — AC for either: launch from the **desktop icon** (the path that failed —
launching from a terminal or the Bridge does not reproduce it), and launching from a terminal or the Bridge does not reproduce it), and
Ardour is still alive 60 s later with `xrun_by.ardour` flat. Ardour is still alive 60 s later with `xrun_by.ardour` flat.
Check whether `perf-audio` does the same `chrt` on Ardour; if so it needs **Epistemic status, so nobody over-trusts this.** The kill mechanism is
the same edit or it will re-arm the bug on the next perf pass. *consistent with every measurement* — SIGKILL with no kernel line, Pulsar
from the SAME launch carrying `rttime=200000` (proving the icon chain
inherits gnome-shell's budget), the hand-launched survivor at `unlimited` —
but it has **not been reproduced**. The AC above IS the reproduction. It is
also non-deterministic by nature: it needs one >200 ms uninterrupted GUI
burst, so earlier icon launches that survived do not falsify it.
The xrun half is measured (`xrun_by.ardour` 0→46, and four prior sessions)
but the FIFO/80 GUI thread as its *cause* is a hypothesis, confounded
tonight by the 65 % cap, `powersave`, a session mid-load, mpv and qjackctl.
What gives it teeth: `/proc/sys/kernel/sched_rt_runtime_us` is at the
default **950000/1000000**, so once the realtime tasks on a runqueue exceed
95 % of a period the kernel throttles **every** RT task on it for the
remaining 50 ms — which is 2.3 quanta at 1024/48k. A GUI thread eating
85 % of a core at FIFO/80 is exactly how `AudioEngine 1` at 83 gets
stalled by something *below* it.
- [ ] **2. The CPU cap, still unpaid** (third session running). Measured now: - [ ] **2. The CPU cap, still unpaid** (third session running). Measured now:
`max_perf_pct=65`, governor `powersave`, EPP `balance_power`, and gig-up's `max_perf_pct=65`, governor `powersave`, EPP `balance_power`, and gig-up's
own line claimed "thermal-mode performance holding" while the gate read own line claimed "thermal-mode performance holding" while the gate read
......
...@@ -56,6 +56,17 @@ SIGXCPU, then SIGKILL, silently. ...@@ -56,6 +56,17 @@ SIGXCPU, then SIGKILL, silently.
- Context that made it bite: `max_perf_pct=65`, governor `powersave`, EPP - Context that made it bite: `max_perf_pct=65`, governor `powersave`, EPP
`balance_power`, while gig-up's own line claimed "thermal-mode performance `balance_power`, while gig-up's own line claimed "thermal-mode performance
holding". Third session that cap has gone unpaid. holding". Third session that cap has gone unpaid.
- **Two writers re-arm it**, which is the other half of the trap:
`parvagues-protect`'s `TARGETS` line and `/usr/local/sbin/perf-audio:565`
(plus 573, for every child at 79) — and `perf-watch` reasserts the second one
by itself on any drift. Fix one and the bug comes back on the next perf pass.
- **What is measured and what is inferred.** The xrun counts and the attribution
are measured. The kill mechanism is consistent with every measurement but was
**not reproduced** — and it cannot be reproduced on demand, because it needs
one >200 ms uninterrupted GUI burst. Mechanism that makes the xrun half
plausible: `sched_rt_runtime_us` sits at the default 950000/1000000, so once
the realtime tasks on a runqueue pass 95 % of a period the kernel throttles
*every* RT task on it for the remaining 50 ms — 2.3 audio quanta.
- Observability verdict: **it worked.** `xrun_by` named the culprit in one query - Observability verdict: **it worked.** `xrun_by` named the culprit in one query
and four sessions of history were already on disk. The gap was never and four sessions of history were already on disk. The gap was never
instrumentation, it was nobody asking the log a question. instrumentation, it was nobody asking the log a question.
......
Markdown is supported
0% or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment