midi-harbor/specs/019-daemon-updates/research.md
James Coleman 57fe7ff53a fix(service): wait for launchd to drop the old job before loading the new
- Update now stopped the daemon and failed with "Bootstrap failed: 5: Input/output error" on macOS, leaving the agent's plist written and no job loaded. `launchctl bootout` returns once the daemon has been sent SIGTERM, and launchd refuses `bootstrap` until the daemon has exited and the job has left the domain.
- Installing or replacing the launchd agent now waits for the booted-out job to leave, for up to 25 seconds, which covers the 20 seconds launchd allows before SIGKILL.
- Starting the service now loads the agent's plist first when launchd does not hold the job, so the window's Start button and `service start` recover a registration left unloaded instead of failing in `kickstart`.
- The stand-in launchctl in the service lifecycle test returns from `bootout` while the daemon is still stopping and refuses `bootstrap` until it has gone, as launchd does, so `service install --start` over a running daemon pins the regression.
2026-10-04 08:43:28 -05:00

5.5 KiB

Research: Daemon Updates

The investigation behind this spec, under its number in the project-wide research log.


R-108: A daemon left running by another build

Status: DONE (2026-10-02). Built as T253.

What happened before. A client called GetServerInfo and refused only another major protocol version. Nothing compared builds, so after an update the old daemon ran on:

Install Updating while the daemon runs
macOS disk image, launchd The old daemon kept running until the next login
Mac App Store Quitting the app stops its daemon, so the updated app started the new one; a daemon left by a window that crashed was attached to and kept
.deb, .rpm Installing starts nothing, so the old daemon kept running
AppImage A new file moved over the old one was used at the daemon's next start
Windows archive No installer; the old daemon kept running

service install --start did not help under systemd or Task Scheduler either: it rewrote the registration and started the service, and starting a service that is running does nothing.

The identifier. midi_harbor_core::BUILD_ID, a UUID set by crates/core/build.rs. It takes MIDI_HARBOR_BUILD_ID from the environment when that is set, and makes one otherwise. A package can hold binaries built separately that must agree: the App Store app and its headless helper, and each architecture of a universal binary. So packaging/macos/build.sh and the release build each choose one identifier for everything they build in a run. A build outside packaging gets a new identifier when the core crate or VERSION changes, which is as often as cargo runs the script again without forcing every build to relink. The daemon reports it as ServerInfo.build_id (7), protocol 1.3; an older daemon leaves it empty, which matches nothing.

One release, several packages. The .deb, .rpm and AppImage of a release hold the same binary and so the same identifier. A user who installs two of them has one build twice, and no daemon to update, so neither window replaces the other's daemon. Comparing the registered path as well was considered and left out: it would restart the daemon every time the other copy was opened, to run the same code.

Why reinstall rather than restart. Restarting runs whatever the service is registered to run, which after a second install is the other copy. Registering this copy and then starting it is what makes the daemon the one the window came with, wherever it is.

Asking instead of acting. The first version replaced the daemon as soon as the window connected. That needed guards against a loop: two windows of different builds each replacing the other's daemon when they reconnected, and a replacement that failed being tried again on every reconnect. The owner asked for a notice with an Update now button instead. Nothing is then restarted unless someone asks, which removes the loop rather than guarding it, and leaves the few seconds without MIDI to be timed by the person at the machine.

What the notice says. "The daemon is outdated" for an older version or one too old to say, "a different build" for the same version, and "newer than this window" for a newer one, where the button reads Use this version: an older program in a newer one's place may not read the configuration it wrote, which is refused with SchemaTooNew. The button is offered only when the service manager says it is running the daemon and the window is on the service's own socket. Otherwise there is no registration that says how the daemon was started, and the notice says to restart it from this copy.

While the daemon is being replaced the old one still answers for a moment. The window does not reconnect until the replacement has finished, or it would put the notice straight back.

Replacing under each service manager. ServiceManager::replace stops the daemon, installs the registration and starts it. Stopping comes first so the service manager stops the process it started. launchd is the exception: installing boots the old job out, which stops its daemon, and stopping it beforehand only has launchd start the old one again in between.

launchd, found in use (2026-10-04). Update now on two Macs running macOS 15.6.1 stopped the daemon and failed with Bootstrap failed: 5: Input/output error, leaving the plist written and no job loaded. launchctl bootout returns once the daemon has been sent SIGTERM, and the job stays in the domain until the daemon exits; bootstrap is refused until then. Reproduced on macOS 27.0.1 with a scratch job that takes 3 s to exit: bootout returned in 0.01 s, print went on answering with state = SIGTERMed and bootstrap failed with error 5 for 3.2 s, then print exited 113 and bootstrap succeeded. The backend now polls print after bootout until the job has left, for up to 25 s, which covers launchd's 20 s before SIGKILL. The Start button then ran kickstart on a job that was not loaded, so start now bootstraps the plist first when launchd does not hold the job. The stand-in launchctl in tests/service_lifecycle.rs behaves the same way, and the test failed against the old backend.

The App Store build has no service. The app starts its helper, or attaches to one already answering. It now asks an attached daemon for its build and, when it is another, asks it to stop and starts its own.

Checked live on Arch Linux under systemd, with two builds of the tree given the identifiers aaaaaaaa-… and bbbbbbbb-…. See T253 for what was seen.