42 KiB
Research: Resilience
Entries from the research log, under their original numbers. Findings marked VERIFIED were proven in this repository; ASSUMED ones rest on the literature.
R-010: Detecting sleep/wake and network changes
Decision: Platform-specific event sources behind a single SystemEvents trait, backed by a
platform-independent heuristic.
| Event | macOS | Linux |
|---|---|---|
| Sleep / wake | IOKit IORegisterForSystemPower |
org.freedesktop.login1.Manager.PrepareForSleep over D-Bus |
| Network change | SCNetworkReachability / nw_path_monitor |
netlink RTMGRP_IPV4_IFADDR / RTMGRP_IPV6_IFADDR |
Status: ASSUMED; moderate risk.
Rationale: These are the events that drive FR-024 and FR-025. The important finding is that
the platform APIs are a convenience layer, not a correctness layer — a machine can suspend
without the API firing, and logind is known not to emit PrepareForSleep(false) when resuming from
hibernation. Relying on them alone would leave exactly the silent-dead-connection failure this
product exists to eliminate.
Therefore: the authoritative recovery mechanism is the per-session liveness check — missed clock-synchronisation responses on RTP-MIDI, and characteristic notification silence on BLE. The platform events are an optimisation that makes recovery near-instant instead of waiting for a timeout. Correctness must not depend on them, and the test suite must prove recovery works with the platform event source disabled.
A third, cheap signal worth adding: a monotonic-versus-wall-clock gap detector. A task that ticks every second and observes that substantially more wall-clock time passed than expected has almost certainly just come back from suspend, on any platform.
Alternatives considered: choreo-power-events — a convenience wrapper over the same APIs;
reasonable, but the trait boundary is needed regardless and the dependency adds little. Polling
interface lists — too slow, too coarse.
R-031: The comment said it silenced; the code did not
Status: VERIFIED (2026-09-20) on macOS over CoreMIDI and Linux over ALSA.
delete_virtual_port carried the comment "Silence before tearing down, so nothing is left
sounding on the far side" directly above the teardown, and the CLI help said "Delete a port,
silencing it first". Neither did. Disabling a port did not either, and stopping the daemon did
not: a held note sounded on whatever received it until something else stopped it, and nothing
else would — the only thing that knew a note was playing had just exited.
EventKind::NotesSilenced had been defined, given a wire name, and never emitted.
Per channel, not all sixteen. Sending a reset on every channel is the usual approach and is
what the network session path already did, but a virtual port is shared: resetting a channel
nothing was playing clears sustain and cuts sound for whatever else is driving that port.
core::sounding keeps a bit per note per channel, updated from what leaves an endpoint, so
silencing touches only the channels that were actually played and an idle port is left alone.
The record is updated from outbound traffic only. A note sounds on whatever received it, so the endpoint that sent it is the one that will have to stop it.
Measured: a held note, then port disable, produced sustain-off, all-notes-off and
all-sound-off on channel 1 alone, on both platforms. Repeating with SIGINT to the daemon
produced the same three, before the socket was removed.
Not covered: a device unplugged mid-note. The endpoint that went away cannot be told anything, and the notes are sounding on its route destinations, which other sources may also be playing. Stopping only what came from the removed source needs per-source attribution, which is route-suspension work (FR-032).
R-033: Nothing was watching the machine
Status: VERIFIED (2026-09-20) by stepping a real machine's wall clock.
The SystemEvents seam existed, a fake existed, Event::Nudge existed in the connection state
machine with a doc comment describing exactly what it was for, and on_retry already reset the
backoff when nudged. Nothing produced a nudge. No backend existed on either platform and nothing
in the daemon consumed the seam, so a machine waking from sleep was invisible.
The number that makes this matter: the backoff caps at thirty seconds. A session that was retrying when the lid closed waits up to thirty seconds after it opens — against SC-003's ten. The mechanism to avoid that was fully built except for the part that notices.
Why the clock gap rather than IOKit and logind first. R-010 already records that the platform power APIs miss suspends. Comparing the monotonic clock against the wall clock catches every suspend on both platforms with no bindings and no privileges, because both platforms stop the monotonic clock while suspended and keep the wall clock running: the difference is how long the machine was away. The platform backends remain worth adding — they report a suspend before it happens, which this cannot — but they are an addition to this, not a replacement for it.
Both failure modes are benign, which is what makes a heuristic acceptable here. A false positive costs one reconnection attempt that was going to happen anyway. A missed detection costs the wait this removes, which is what happens today. Neither can lose MIDI.
Nudges are ignored by sessions that are connected or switched off, so a hint arriving while everything is fine cannot turn into a dropout.
Measured: stepping the wall clock forward twenty seconds on the Linux machine produced
away_seconds=19 in the log and a system_resumed event, through the whole chain from detector
to event log.
The network half, by the same reasoning. NetworkChanged is produced by comparing the set of
addresses this machine answers on, loopback excluded. An address appearing or disappearing is how
a network comes back without anything announcing it — a cable plugged in, a laptop joining a
different network, a lease renewed on a different address — and a session bound to the old address
never recovers by waiting. if-addrs was already in the dependency tree through mdns-sd, so
this costs nothing in the headless build.
The platform-native sources are still worth adding, for the same reason as the power ones: they report the change the instant it happens rather than within a second, and they can say what changed. They are an addition, not a replacement (T097, T098).
Measured: adding an address to a throwaway interface on the Linux machine produced
addresses=30 and a system_resumed event; removing it produced addresses=28 and another. The
event detail names which change prompted the reconnection, because a reconnection that happens on
its own reads as a spontaneous dropout otherwise.
R-034: A suspended session looks exactly like a working one
Status: VERIFIED (2026-09-20) between the macOS and Linux machines, measured both ways.
R-033 framed the wake case as a session waiting out a thirty-second backoff. That was the wrong mechanism, and the nudge built for it did nothing at all in the case it was named after.
After a suspend, a session is still in Connected. Nothing ran while the machine was away, so
nothing failed and nothing was scheduled. The nudge only acted on sessions that were retrying, so
it skipped the one case that mattered. What actually notices is LIVENESS_TIMEOUT, which is
thirty-five seconds — more than twice SC-004's budget, and the backoff never enters into it.
Thirty-five seconds is right for an ordinary quiet session: a link that is merely idle is not a link that is broken, and a shorter timeout would drop working sessions. What was missing is that a hint changes the question. After a wake or an address change the link is suspect, and the right move is to ask the peer now and give up in seconds rather than in half a minute.
Session::probe sends a clock exchange immediately and sets a two-second deadline. An answer
clears it, so a working link survives any number of hints. Repeated hints do not push the deadline
out, or a machine reporting changes steadily would never let a probe conclude — which is not
hypothetical: adding one interface produced three address changes in three seconds.
Measured, by dropping the session's UDP between the two machines so the link died silently:
| time to notice | |
|---|---|
| with a hint to prompt a probe | 3.8 s |
| with nothing to prompt one | 37.2 s |
The second number is the behaviour every release before this one had, and it is what SC-004's fifteen seconds was being missed by.
R-035: A thirty-second retry ceiling cannot meet a ten-second promise
Status: VERIFIED (2026-09-20) by simulation over the real backoff and measured between the two machines.
Writing the test for SC-003 was what found this. An interruption that leaves this machine's own addresses alone — an upstream router rebooting, a switch power-cycling — reports nothing when it ends. The address watch cannot see it and the peer cannot tell us. The only thing that finds out is the next scheduled retry, so the ceiling on that delay is the recovery time.
The ceiling was thirty seconds, with full jitter. Simulating a hundred interruptions over the real backoff: 53 of 100 took longer than ten seconds, some over twenty-five. SC-003 asks for at most one in a hundred. It was never close.
Thirty seconds is a sensible default for a retry that costs something. A network session's retry
is three small packets, and waiting half a minute between them saves nothing worth having while
being most of the time a user spends wondering why their keyboard stopped working.
BackoffPolicy::responsive caps at five seconds and is what network sessions use; the default
stays for everything else.
Measured on the two machines, blocking the session's UDP without touching any address: the session gave up, retried through a twenty-second outage, and reconnected 2.8 seconds after the block was lifted.
The division of evidence is deliberate. The distribution — "at least 99 of 100" — is a property of
the jitter and the ceiling, which a simulation over the real Backoff can assert every run. That
one interruption recovered in 2.8 seconds is evidence the mechanism works end to end, not evidence
about the ninety-nine.
R-044: A flapping link was retried four times a second, forever
Status: FIXED (2026-09-21) in the state machine, with the delays tested against an injected clock.
The spec's rapid connect/disconnect edge case asks two things: the backoff keeps a flapping link from consuming resources, and the user can see that the link is unstable. The state machine did neither. It reset the backoff on every successful connection, so a link that came up and dropped at once went straight back to the shortest delay. Its eight retries were spaced at 250 ms, 250 ms, 250 ms and so on, which is what the regression test prints when the old reset is put back. Nothing distinguished it from a link that had simply failed once.
A connection now resets the backoff only once it has stayed up for STABLE_AFTER (10 s). Until
then the backoff carries on growing, so the same eight flaps end with retries 16 s or more apart.
After UNSTABLE_AFTER (3) brief connections in a row the state reports the link as unstable,
until a connection lasts. ConnectionState.unstable carries this over the wire (field 6, added
within midiharbor.v1). The CLI shows connected (unstable), and the GUI says it keeps dropping
within seconds of connecting.
Two further findings:
- Bluetooth reconnected on sight. A remembered device was reconnected every time it was heard, whatever its backoff said, so the damping above never reached it. It now waits out the backoff. Simply ignoring an early sighting would be worse than no damping, because the radio may not report the device again. So an early sighting schedules the reconnect for when the backoff ends. The test flaps a device until it is unstable, then brings it back once, and requires that single sighting to reconnect it. Dropping early sightings instead fails it.
- The session supervisor had a second backoff that decided nothing. It was reset on connect, on nudge and on success, and never asked for a delay. Retry timing always came from the state machine. It was removed.
Not covered: hardware that flaps by being unplugged and replugged gets a fresh state on each arrival, because the operating system decides when it comes back, so there is nothing to space out and no history to call it unstable by.
R-045: "No network" and "the peer is not answering" were the same message
Status: VERIFIED (2026-09-21) on Linux, with one daemon inside unshare -n, a network
namespace with no network at all, and one outside it.
The spec's no-network edge case asks that sessions report waiting for a network rather than
repeatedly reporting connection failures. Every unanswered invitation was reported as
network unreachable, with a warning for each retry. That happened with no network at all, and
equally on a working network where the peer was simply not running.
Deciding that there is no network. Counting the machine's addresses does not work: an
interface that is up carries an IPv6 link-local address with no network behind it, and on macOS
tunnel interfaces carry one even with Wi-Fi off. What is reliable is the kernel's own answer. A
send with no route is refused at once with NetworkUnreachable, HostUnreachable or
AddrNotAvailable, before anything reaches the wire. The session marks itself waiting for a
network when that happens and clears the mark on the next send that goes out.
ConnectionState.waiting_for_network carries it (field 7, added within midiharbor.v1).
| Before | After | |
|---|---|---|
No network (unshare -n) |
retrying, network unreachable, a warning per retry |
waiting for network, logged once |
| Network up, peer not running | network unreachable, a warning per retry |
peer not responding, one info line |
The PeerTimeout message changed from "peer stopped responding" to "peer not responding", which
also fits a peer that never answered. The slug, which is the contract, is unchanged. A lost
connection still warns every time; only repeated failed attempts went quiet.
R-050: One endpoint being retried stopped MIDI everywhere
Status: FIXED (2026-09-21), for T071a (FR-029).
FR-029 requires healthy connections to carry on while another is failing or retrying. Network
sessions and Bluetooth links always ran on tasks of their own. Virtual ports and hardware that
failed to open were retried by one loop in supervisor.rs, and that loop broke FR-029 three ways:
- It held the daemon's lock through the platform open. Every route takes that lock to find its destinations, so a device held by another application, retried every few seconds, stopped all MIDI in the daemon for as long as each attempt took to answer. With opens that took two seconds, a healthy route between two other ports stalled for 1.98 s. The lock is now taken to claim the attempt and to record its outcome, and released in between.
- It ran the open on a runtime worker. A backend's open waits for its platform thread to
answer. Waiting on a worker held up whatever task tokio had queued behind it there, which no
idle worker can take over. With the lock released but the open still on a worker, the
daemon's own tasks and the test's alike stalled for the full two seconds. Opens now run through
spawn_blocking. - It retried one endpoint after another. Three devices slow to answer took 4.7 s to come back, against 1.5 s for one. A fault in one attempt would also have ended retrying for every endpoint. Each attempt now runs on its own task.
Releasing the lock opens a window: the endpoint can be switched off, deleted or reopened while the platform answers. Switching it off or deleting it removes its runtime, and switching it back on starts a new one. The outcome is therefore recorded only onto the same attempt's runtime, still connecting with nothing open. Otherwise whatever the attempt opened is closed again.
crates/daemon/tests/isolation.rs covers all three defects and the window:
- a healthy route stays under 250 ms throughout two seconds of slow retries;
- a device switched off mid-retry is not left open;
- three slow devices come back in parallel.
The fake platform gained delay_opens to make opens slow. Each test fails with its fix reverted.
One case is not induced: a re-enable whose own open finishes before the retry's. A uniform delay
cannot order the two that way, so the guard against it is reasoned rather than tested.
Two other opens are unchanged:
reconcilestill opens virtual ports under the lock. It runs when the configuration changes, which is a user acting rather than an endpoint failing.- Hardware plugged in is opened outside the lock but on a worker.
Both happen once per change rather than every backoff step, and a real backend answers in milliseconds. They matter only if an open ever hangs.
R-051: Releasing held notes before the machine sleeps
Status: VERIFIED (2026-09-21), for T097 and T098, with a real suspend on each platform in R-070 and R-072.
R-033 added the polled watcher, which finds a suspend only once the machine is back. That is enough to reconnect (FR-024), but too late for one thing: a note held on another machine. When this machine sleeps, the far side hears nothing more. It releases the note only when its own liveness check gives up, more than half a minute later. Only the platform says a suspend is coming, and both platforms will wait for a process that asks them to:
| Announcement | Hold | |
|---|---|---|
| Linux | logind PrepareForSleep(true) |
a delay inhibitor lock, released by closing its descriptor, bounded by InhibitDelayMaxSec (5 s by default) |
| macOS | IOKit kIOMessageSystemWillSleep |
answering with IOAllowPowerChange; an unanswered message holds the machine for 30 s |
On a suspend still to come, the daemon silences every route, network sessions and Bluetooth
links included. It then calls SystemEvents::ready_for_sleep, and the backend lets the suspend
go. Each backend also lets it go after READY_BOUND (3 s) on its own: a stalled daemon must not
keep a machine awake. The polled watcher reports a suspend only paired with its resume, after the
fact, and the daemon ignores a suspend that its resume already follows. Releasing notes on waking
could cut one someone has just started playing.
Both backends run beside the polled watcher (CombinedSystemEvents), never instead of it, since
R-010 records that both platforms miss suspends. They report sleep and wake only. Network changes
stay with the polled address watch, which finds them within a second. A native source would save
that second and nothing more.
Measured on Debian 12, with the real binary and a scratch configuration:
systemd-inhibit --listshowsMidi Harbor … sleep … delaywhile the daemon runs, and nothing after it exits.- A forged
PrepareForSleep(true)sent withdbus-sendis ignored: only logind's own bus name is believed. - With that sender check removed in a throwaway copy, the same forged signal went through the
whole chain:
- the daemon logged
releasing held notes; - the lock was gone within two seconds;
PrepareForSleep(false)took it again.
- the daemon logged
- The check restored, the forgery is ignored once more.
On macOS, registers_for_power_notifications confirms that IOKit accepts the registration,
which is the step that fails if the declarations are wrong. Answering the messages has only been
reasoned through: every kIOMessageSystemWillSleep path reaches IOAllowPowerChange.
crates/daemon/tests/suspend.rs covers the daemon's half. A note held through a loopback
session is released at the far synth on a pending suspend, and the platform is told it may
proceed. A suspend already over releases nothing. Each fails with its rule reverted.
R-067: The 24-hour soak, and a session between the two machines
Status: DONE (2026-09-22).
The soak (T148, SC-007) ran on the Linux desktop for a full day, from 0206fa6: two daemons in
one process joined by a session, with faults injected in turn. It finished with
done: 24278 faults over 86400.134791996s, every one recovered, slowest 2.558007105s,
11511276 messages played
Every fault recovered, and the slowest took 2.56 s against SC-003's ten seconds. The run does not cover the session work done the same day, which landed after it started.
A session between this Mac and the desktop, over the LAN, with today's build on both:
- A 3002-byte dump and a note behind it went from the Mac's port to the desktop's, arriving whole and in order, and then the same in the other direction, checked byte for byte (R-066).
- The Mac disconnecting left both sides listening. The desktop did not invite the Mac back, and its history reads "MacLink ended its session with PCLink; PCLink is listening again".
- On the desktop, nothing was remembered but the kernel's own Midi Through port: the
aseqdumpand BlueZ clients that had opened ports during the day left no entry behind.
A second soak (2026-09-23 to 24) covered the session and routing work since. The first
attempt, from before c4f3c1a, failed in hour 22: at cycle 22,579 a keyboard replug was not
recovered within ten seconds, and the run printed nothing that said why. c4f3c1a made an
unrecovered fault print each side's session phase, attempt and peer, every route, the last 30
events and the keyboard's handle. The run from c4f3c1a on the Arch VM then finished clean:
done: 24654 faults over 86402.498228452s, every one recovered, slowest 3.215248496s,
10316356 messages played
The hour-22 failure did not recur, so its cause is unknown. If it comes back, the diagnostics will say what state it was in.
R-069: Roaming between two access points
Status: DONE (2026-09-22), for T154 scenario 2. The session survived one roam and reconnected a second after the other. No note was left sounding. The history blamed the peer for a network this machine had lost, which is fixed.
The owner has two access points with the same network name and authentication. A session ran
from this Mac, on Wi-Fi, to the Linux desktop on the LAN, today's build at each end. The desktop
routed the session to a virtual port whose output aconnect fed back into its input, so whatever
the Mac sent came back. A CoreMIDI client on the Mac sent a numbered message every 50 ms and
logged when each returned. A CoreWLAN watcher logged the Wi-Fi channel, since macOS hides the
access point's identity from unprivileged tools. Each roam was forced by restarting the access
point the Mac was on. The Mac kept 192.0.2.10 throughout, so neither roam tested a change of
address.
First roam, pitch bend. Messages sent from 14:40:40.95 to 14:40:46.45 were lost, 103 of them, and the Mac was on the other access point (channel 36 to 157) at 14:40:46.65. The session never dropped, and messages came back from the moment the Mac rejoined. The round trip, 7 ms before, ran up to 180 ms for about half a minute on the new access point and then settled. The counters put almost all the loss on the Mac's side: 4,565 sent, 4,467 received by the desktop. Every packet in this stream carried a newer pitch bend than its journal, so recovery correctly applied nothing, and this pass shows only that the session held.
Second roam, notes. Each 50 ms tick turned the previous note off and the next one on. Messages stopped coming back at 14:48:02.86. The Mac left its access point at 14:48:13, lost its addresses at 14:48:16, and joined on channel 40 at 14:48:35. The history read:
14:48:16 system_resumed this machine's addresses changed; reconnecting
14:48:18 endpoint_state_changed MacLink lost its connection to PCLink: peer not responding
14:48:36 system_resumed this machine's addresses changed; reconnecting
14:48:37 endpoint_state_changed MacLink connected to PCLink after 3 failed attempts
The session reconnected 1.0 s after the addresses came back, well inside SC-003's ten seconds. Dropping the session released the notes sounding through it. At the end, of 9,795 notes played and 9,198 heard back, none was left sounding.
The lost connection blamed the peer. At 14:48:15.6 the session logged that it had no network to reach the peer by. Three seconds later its check on the peer timed out and was recorded as "peer not responding". A timeout was always reported as the peer's; only unanswered invitations took the missing network into account. A check that timed out with nothing able to leave the machine is now "network unreachable" as well.
Not covered: a roam that changes the address, which needs two networks, and messages the network loses during an outage. Those are only as recoverable as the journal makes them, as in SC-006.
R-070: Sleep and wake, and a session that carried nothing
Status: DONE (2026-09-22), for T154 scenario 2. A session from this Mac to the Linux desktop comes back within half a second of the Mac waking, and no note is left sounding. The first run found a session that said it was connected and carried nothing, and the desktop kept waking the sleeping Mac. Both are fixed.
The setup was R-069's: the Mac, on Wi-Fi and mains power, invited the desktop, which sent back
whatever arrived. A note played every 50 ms, each turning the previous one off. The owner slept
the Mac from the Apple menu and woke it by hand. pmset -g log gave the Mac's own sleep and wake
times.
The Mac did not stay asleep. Each time, about 32 s after the Mac went to sleep, the desktop's
liveness check gave up on it and the desktop invited it again. Within a second the Mac woke
briefly, a "DarkWake" that pmset puts down to wifibt. The session reconnected, the Mac slept
again about 45 s later, and the cycle repeated. The first run woke five times this way, between
15:08:10 and 15:13:27, and the second once, at 15:30:46. Every one came a second after the
desktop's invitation. So a Midi Harbor peer keeps a sleeping Mac on mains power awake for about
45 s in every 80. The desktop keeps inviting because a session whose peer went quiet chases
it, whichever side invited: only a peer that says goodbye is let go.
A session that said it was connected carried nothing. After the owner woke the Mac at
15:14:03, both ends read connected. The Mac sent a note every 50 ms, and the desktop's received
count stayed at 3,706. Clock exchanges went both ways, so neither side timed out. The Mac's
packets carried sequence 0x72af and a journal checkpoint of 0xab65, which is ahead of
anything that session had sent. At the last dark wake, the desktop's fresh session had taken in
MIDI still coming from the Mac's old session, whose sequence numbers were near 0xab65, because
a session took MIDI from any source. The Mac then began a new session from a random sequence
14,500 packets behind that. The desktop dropped every one of those packets as late, and would
have gone on dropping them for about twelve minutes. Its feedback acknowledging 0xab65 also
trimmed the Mac's new journal. A session now takes MIDI and feedback only from the source agreed
in the handshake.
On the fix, the second run slept at 15:30:13. The desktop re-invited at 15:30:45, and the Mac dark-woke and reconnected at 15:30:46. MIDI was back 31.5 s after it stopped, and it carried on through the owner's full wake at 15:31:05. The only note sounding afterwards was the one playing. SC-004 asks for a reconnect within 15 s of waking, and both runs reconnected within a second of each wake.
Saying goodbye before sleep. The owner chose to have a machine end its sessions as it goes to sleep, so the peer stops inviting it. The first try sent an ordinary goodbye and reconnected after five seconds awake without a wake notice. It failed twice on the Mac:
- macOS went on running for five seconds after the notice of sleep, and a slow Bluetooth acknowledgement for another one and a half. The fallback counted that as time awake and reconnected just as the Mac went down, so the desktop woke it 35 s later. The fallback is now sixty seconds awake. Normally the wake notice reconnects at once.
- When the desktop had made the connection, an ordinary goodbye is what a restarting peer sends, so the desktop invited again 0.3 s later. The Mac, not yet asleep, accepted. Who made the connection also drifts: whichever side re-invites after a timeout becomes the one that made it.
The goodbye before sleep now carries SLEP in its token field, which other implementations
ignore on a goodbye. A Midi Harbor peer that receives it goes back to listening, whichever side
made the connection, and the sleeping machine reconnects on waking. Checked with the desktop
having made the connection: the Mac slept at 17:23:43 for 62 s. pmset shows no wake in that
time, and the desktop's history reads "MacLink ended its session with PCLink; PCLink is listening
again". The owner woke the Mac at 17:24:45, and the session was up 0.4 s later. Apple and other
implementations see an ordinary goodbye and may still invite a sleeping Mac.
R-072: Sleep on Linux, and NetworkManager going first
Status: DONE (2026-09-22). A Linux machine going to sleep now tells its peer, as the Mac does (R-070). Before, NetworkManager had the network down before the daemon looked.
An Arch Linux virtual machine slept through logind with
systemctl suspend, into s2idle. An RTC alarm did not wake it from there, and a key press through
virsh send-key did. A Midi Harbor session on the VM was connected to one on the Mac, by address,
and tcpdump on the Mac recorded what reached it.
What happened first. logind announced the suspend at 18:38:58.917, and NetworkManager, reacting
to the same signal, took eth0 down 0.2 ms later, finishing in 31 ms. The daemon holds a delay
lock, so logind waited for it, but it looked for events only once a second. It saw the suspend at
18:38:59.19 and found no network to send the goodbye by, so the Mac heard nothing. The Mac timed
out 32 s later ("peer not responding") and invited the sleeping machine until it came back. The
notes released before sleep went nowhere for the same reason, so on a machine running
NetworkManager, notes held on the peer sounded until its liveness check gave up.
Worse, NetworkManager's teardown arrived in the same batch as an address change, and the daemon read that as a reason to reconnect. The session just ended for sleep was reconnected at once, on the way down.
The fix. A backend that reports a suspend now signals the daemon, which acts on it at once rather than at its next look. The watcher also remembers that the machine is going to sleep and ignores address changes until it wakes, or until it has been awake as long as the sessions wait (R-070). On the next suspend the daemon ended its session at 18:46:58.7111, 0.2 ms before NetworkManager logged the request. The Mac, which had made the connection, recorded "VMSide ended its session with MacSide; MacSide is listening again", and nothing reached the VM until it woke 50 s later. It reconnected 0.8 s after waking.
This is a race the daemon now wins by a fraction of a millisecond, not an ordering anything guarantees: NetworkManager does not wait for other holders of the delay lock. A slower machine may still lose it now and then, and the peer then falls back to its liveness check as before.
R-074: Restoring controller state on a recovered link
Status: DONE (2026-09-22), for T155 (FR-027, constitution Principle I). A converge pass found that the connection state machine asked for a restoration on every recovery and nothing acted on it. A fresh session starts with an empty recovery journal, so a volume moved while a link was down stayed stale at the receiver, and a Bluetooth synth switched off and on came back at its defaults.
midi_harbor_core::controls::Controls keeps, per channel, the last value of every controller
below the channel-mode range, the program, the pitch bend and the channel pressure an endpoint was
sent. On recovery it produces, per channel, bank select, then the program it qualifies, then the
other controllers in order, then pitch bend and pressure.
Decisions:
- Where it is kept. A network session's supervisor records everything routed into it, and keeps the record across attempts, since routes to a session stay live while it reconnects. A control moved during an outage is therefore restored. A Bluetooth link's routes are suspended while it is down, so its record holds what was last delivered: a device that lost its settings gets them back, and a change made during the outage does not reach it.
- What is left out. Notes, which FR-026 silences and which must never be replayed; the channel-mode controllers (120 and up), which are commands, not state; and data entry, increment, decrement and the RPN and NRPN selectors (6, 38, 96 to 101), which only mean something in sequence. Replayed in numeric order, data entry would land on whichever parameter the later selector chose. A stale pitch-bend range is the lesser harm than a wrong one.
- Reset all controllers clears the channel's record except its program, which it does not reset.
Tests: crates/core/src/controls.rs for the rules, crates/daemon/tests/restoration.rs for a
fader moved while a session was down, and a Bluetooth test for a device that drops and returns.
Each fails with the restoration switched off.
R-079: Hardware on real Macs, and a MIDI server that dies
Status: CLOSED (2026-09-25). Recovery on macOS is built and checked (T192). Linux has nothing to recover from, and its checks with the pad controller turned up four smaller faults, all fixed (T193 to T195). Replacement under an installed launchd service was checked on 2026-09-25, with the owner's permission, and works.
Two Macs, then one Linux machine, were checked with real USB MIDI hardware.
A Mac running a presentation program, with a generic USB MIDI interface (maker "Generic", model
"USB Midi ", with a trailing space) that has nothing cabled to it. A scratch daemon with its own
config and socket, built without Bluetooth so no permission prompt reached that screen, found the
interface, read its maker and model, bound routes to it by name despite the trailing space,
refused a two-route loop, carried 8 messages out to it, kept its virtual ports' CoreMIDI ids across
a restart, and paused and resumed its routes across device disable and device enable. Two
CLI labels were wrong there: the program's Apple network session read "physical", and a
switched-off device read "attached". Both are fixed (they read "provided" and "disabled").
The development Mac, with a USB pad controller. Pads routed to a virtual port arrived through the route in 0.26 ms median, 6.3 ms worst, with every note-on and note-off accounted for. A first count seemed to lose note-offs; the listener read only the first message of each CoreMIDI packet, and a corrected listener saw 5 of 5.
The first unplug crashed Apple's MIDIServer (SIGSEGV in MIDIObject::GetIntegerProperty,
under MIDIGetNumberOfSources and ObjectTreeCache::GetObjectTree, serving a client's request
for the object tree while the device was being torn down). The daemon re-reads the device list
the moment a removal is reported, so it may have been the client that asked; the notification
logger that also enumerated inside its callback did not crash the server on a later unplug, so
this is a race, not a certainty. launchd restarted the server, and the daemon never recovered:
the device stayed absent after it was plugged back in, and creating a port failed with
OSStatus -50.
What a client sees when MIDIServer dies, measured by killing it (killall MIDIServer, which
runs as the user):
- No notification of any kind.
- Every call that reaches the server fails from then on: a property of the client's own endpoint and creating a port return -50, finding its own endpoint by unique id returns -10842.
- Sending on its own virtual source returns 0 and goes nowhere.
- Cached counts, such as the number of sources, stay at their old values.
- Creating a new client in the same process also returns -50, and still did 150 s later with a
new
MIDIServerrunning and serving other processes.
So nothing inside the process can recover. The daemon has to notice by asking, since nothing tells it, and has to be replaced by a new process. The owner accepted recovery taking longer than the usual five seconds for this case, as long as it comes back without anyone stepping in.
The recovery (T192). Every two seconds the CoreMIDI thread creates and drops a private output
port; a failure means the server is gone. The daemon then stops serving, silences what it can and
ends its network sessions, and replaces itself with the same executable and arguments through
exec, so the process identifier launchd or systemd knows does not change. The new process
reads the configuration, recreates the ports with their pinned identifiers, reopens the hardware
and routes, and records a warning saying why it started. If the server is not answering yet, it
waits up to thirty seconds for it rather than failing. Killing MIDIServer under a scratch
daemon with the pad controller routed:
- The probe noticed 2.8 s after the kill, and the new process was serving 0.4 s later.
- The virtual ports came back with the same CoreMIDI identifiers, and a note crossed a routed pair of them.
- The owner's pads came through the route, 5 note-ons and 5 note-offs of 5, and after an unplug and replug, 11 of 11. That unplug did not crash the server; the crash is a race, which is why killing the server is how this is tested.
- The window showed "cannot be reached" while the daemon was down and filled in again by itself.
Stopping with the window open never finished. Found while checking the window: stopping
waited for every connection to close, and a window keeps streams open that never close by
themselves. A daemon asked to stop, by Ctrl-C, by service stop or to be replaced, kept running
until it was killed, and a kill skips silencing held notes. Stopping now waits two seconds for
requests already being answered and then closes the remaining connections. With the window open,
a stop request now ends the daemon in 2 s.
Under launchd. The app bundle was built from the current code, copied to ~/Applications,
and its binary ran service install --start on the development Mac, with a fresh configuration
holding one port, "Launchd Check". killall MIDIServer at 18:45:47 UTC. The probe reported the
server gone at 18:45:51.46, the daemon logged that it was restarting to reach it, and the new
process was listening at 18:45:51.96 with the port open again. launchd listed the same process
identifier, 67582, and a last exit status of 0 before and after: it saw no exit and started
nothing. The daemon started exactly twice in its log, and the history read "restarted because the
MIDI server stopped answering; ports, hardware and routes were set up again". The service,
configuration, log and app copy were removed afterwards.
Linux, with the pad controller. On an EndeavourOS machine (kernel 6.18), a headless build from the Arch VM ran as a scratch daemon with its own config and socket. Nothing was built on that machine and no kernel module was touched.
- Pads routed to a virtual port arrived identical and in order, 9 of 9 notes, under 1 ms behind a listener on the device itself (shell timestamps, which resolve no finer).
- Seven unplugs. Each time the device was reopened about 10 s after it went away, the route went to waiting and back, and notes came through afterwards.
- A pad held through an unplug was silenced 1 ms after the daemon saw the removal.
- There is no server to lose. The ALSA sequencer is part of the kernel, not a process that can
crash and be restarted under a client. Unloading
snd-seqis the nearest equivalent, and it cannot be done while PipeWire holds it, so it was not tried.
Four faults turned up there:
Midi Throughread "physical". A client counted as hardware when its number was below 128.Midi Throughis a kernel client numbered among the sound cards, with no card. Belonging to a card is now what makes a client hardware (T193).- Other programs' ports were described as hardware.
aseqdumpwas logged as "opened attached hardware", and a routed application port that closed would have been recorded as "unplugged". Those now read "provided port", "went away" and "is back" (T193). - Every replug zeroed the device's counters. The pads read "received 0" beside a route from them that had carried 58 messages. A device's counters now last while it is unplugged and carry on when it returns (T194).
- No note-off was sent for a held note. Silencing sent sustain-off, all-notes-off and all-sound-off, while the comment on it said a synth ignoring all-notes-off would still honour explicit note-offs. Nothing sent them, though every held note is tracked. Each held note now gets its own note-off, after the pedal release. The same check found the broad resets breaking the rule above that only the notes the stopped route carried may be stopped: two controllers on the drum channel into one sampler, and unplugging one cut the other's held note. A channel another route still has notes sounding on, at the same destination and MIDI Out, now gets only the stopped route's note-offs. A channel nobody else is playing still gets the full reset, which is what stops a drone or an envelope that ignores note-off (T195). Measured: a held pad, then the cable pulled, and the virtual port received sustain-off, note-off 82, all-notes-off and all-sound-off, the received count carried from 15 to 17, and the route agreed at 17.