midi-harbor/specs/019-daemon-updates/research.md
James Coleman 57fe7ff53a fix(service): wait for launchd to drop the old job before loading the new
- Update now stopped the daemon and failed with "Bootstrap failed: 5: Input/output error" on macOS, leaving the agent's plist written and no job loaded. `launchctl bootout` returns once the daemon has been sent SIGTERM, and launchd refuses `bootstrap` until the daemon has exited and the job has left the domain.
- Installing or replacing the launchd agent now waits for the booted-out job to leave, for up to 25 seconds, which covers the 20 seconds launchd allows before SIGKILL.
- Starting the service now loads the agent's plist first when launchd does not hold the job, so the window's Start button and `service start` recover a registration left unloaded instead of failing in `kickstart`.
- The stand-in launchctl in the service lifecycle test returns from `bootout` while the daemon is still stopping and refuses `bootstrap` until it has gone, as launchd does, so `service install --start` over a running daemon pins the regression.
2026-10-04 08:43:28 -05:00

84 lines
5.5 KiB
Markdown

# Research: Daemon Updates
The investigation behind this spec, under its number in the project-wide research log.
---
## R-108: A daemon left running by another build
**Status**: **DONE** (2026-10-02). Built as T253.
**What happened before.** A client called `GetServerInfo` and refused only another major
protocol version. Nothing compared builds, so after an update the old daemon ran on:
| Install | Updating while the daemon runs |
|---|---|
| macOS disk image, launchd | The old daemon kept running until the next login |
| Mac App Store | Quitting the app stops its daemon, so the updated app started the new one; a daemon left by a window that crashed was attached to and kept |
| `.deb`, `.rpm` | Installing starts nothing, so the old daemon kept running |
| AppImage | A new file moved over the old one was used at the daemon's next start |
| Windows archive | No installer; the old daemon kept running |
`service install --start` did not help under systemd or Task Scheduler either: it rewrote the
registration and started the service, and starting a service that is running does nothing.
**The identifier.** `midi_harbor_core::BUILD_ID`, a UUID set by `crates/core/build.rs`. It takes
`MIDI_HARBOR_BUILD_ID` from the environment when that is set, and makes one otherwise. A package
can hold binaries built separately that must agree: the App Store app and its headless helper,
and each architecture of a universal binary. So `packaging/macos/build.sh` and the release build
each choose one identifier for everything they build in a run. A build outside packaging gets a
new identifier when the core crate or `VERSION` changes, which is as often as cargo runs the
script again without forcing every build to relink. The daemon reports it as
`ServerInfo.build_id` (7), protocol 1.3; an older daemon leaves it empty, which matches nothing.
**One release, several packages.** The `.deb`, `.rpm` and AppImage of a release hold the same
binary and so the same identifier. A user who installs two of them has one build twice, and no
daemon to update, so neither window replaces the other's daemon. Comparing the registered path as
well was considered and left out: it would restart the daemon every time the other copy was
opened, to run the same code.
**Why reinstall rather than restart.** Restarting runs whatever the service is registered to
run, which after a second install is the other copy. Registering this copy and then starting it
is what makes the daemon the one the window came with, wherever it is.
**Asking instead of acting.** The first version replaced the daemon as soon as the window
connected. That needed guards against a loop: two windows of different builds each replacing
the other's daemon when they reconnected, and a replacement that failed being tried again on
every reconnect. The owner asked for a notice with an Update now button instead. Nothing is then
restarted unless someone asks, which removes the loop rather than guarding it, and leaves the
few seconds without MIDI to be timed by the person at the machine.
**What the notice says.** "The daemon is outdated" for an older version or one too old to say,
"a different build" for the same version, and "newer than this window" for a newer one, where the
button reads Use this version: an older program in a newer one's place may not read the
configuration it wrote, which is refused with `SchemaTooNew`. The button is offered only when the
service manager says it is running the daemon and the window is on the service's own socket.
Otherwise there is no registration that says how the daemon was started, and the notice says to
restart it from this copy.
**While the daemon is being replaced** the old one still answers for a moment. The window does
not reconnect until the replacement has finished, or it would put the notice straight back.
**Replacing under each service manager.** `ServiceManager::replace` stops the daemon, installs
the registration and starts it. Stopping comes first so the service manager stops the process it
started. launchd is the exception: installing boots the old job out, which stops its daemon, and
stopping it beforehand only has launchd start the old one again in between.
**launchd, found in use (2026-10-04).** Update now on two Macs running macOS 15.6.1 stopped the
daemon and failed with `Bootstrap failed: 5: Input/output error`, leaving the plist written and
no job loaded. `launchctl bootout` returns once the daemon has been sent SIGTERM, and the job
stays in the domain until the daemon exits; `bootstrap` is refused until then. Reproduced on
macOS 27.0.1 with a scratch job that takes 3 s to exit: `bootout` returned in 0.01 s, `print`
went on answering with `state = SIGTERMed` and `bootstrap` failed with error 5 for 3.2 s, then
`print` exited 113 and `bootstrap` succeeded. The backend now polls `print` after `bootout` until
the job has left, for up to 25 s, which covers launchd's 20 s before SIGKILL. The Start button
then ran `kickstart` on a job that was not loaded, so `start` now bootstraps the plist first when
launchd does not hold the job. The stand-in `launchctl` in `tests/service_lifecycle.rs` behaves
the same way, and the test failed against the old backend.
**The App Store build** has no service. The app starts its helper, or attaches to one already
answering. It now asks an attached daemon for its build and, when it is another, asks it to stop
and starts its own.
**Checked live** on Arch Linux under systemd, with two builds of the tree given the identifiers
`aaaaaaaa-…` and `bbbbbbbb-…`. See T253 for what was seen.