- Update now stopped the daemon and failed with "Bootstrap failed: 5: Input/output error" on macOS, leaving the agent's plist written and no job loaded. `launchctl bootout` returns once the daemon has been sent SIGTERM, and launchd refuses `bootstrap` until the daemon has exited and the job has left the domain. - Installing or replacing the launchd agent now waits for the booted-out job to leave, for up to 25 seconds, which covers the 20 seconds launchd allows before SIGKILL. - Starting the service now loads the agent's plist first when launchd does not hold the job, so the window's Start button and `service start` recover a registration left unloaded instead of failing in `kickstart`. - The stand-in launchctl in the service lifecycle test returns from `bootout` while the daemon is still stopping and refuses `bootstrap` until it has gone, as launchd does, so `service install --start` over a running daemon pins the regression.
84 lines
5.5 KiB
Markdown
84 lines
5.5 KiB
Markdown
# Research: Daemon Updates
|
|
|
|
The investigation behind this spec, under its number in the project-wide research log.
|
|
|
|
---
|
|
|
|
## R-108: A daemon left running by another build
|
|
|
|
**Status**: **DONE** (2026-10-02). Built as T253.
|
|
|
|
**What happened before.** A client called `GetServerInfo` and refused only another major
|
|
protocol version. Nothing compared builds, so after an update the old daemon ran on:
|
|
|
|
| Install | Updating while the daemon runs |
|
|
|---|---|
|
|
| macOS disk image, launchd | The old daemon kept running until the next login |
|
|
| Mac App Store | Quitting the app stops its daemon, so the updated app started the new one; a daemon left by a window that crashed was attached to and kept |
|
|
| `.deb`, `.rpm` | Installing starts nothing, so the old daemon kept running |
|
|
| AppImage | A new file moved over the old one was used at the daemon's next start |
|
|
| Windows archive | No installer; the old daemon kept running |
|
|
|
|
`service install --start` did not help under systemd or Task Scheduler either: it rewrote the
|
|
registration and started the service, and starting a service that is running does nothing.
|
|
|
|
**The identifier.** `midi_harbor_core::BUILD_ID`, a UUID set by `crates/core/build.rs`. It takes
|
|
`MIDI_HARBOR_BUILD_ID` from the environment when that is set, and makes one otherwise. A package
|
|
can hold binaries built separately that must agree: the App Store app and its headless helper,
|
|
and each architecture of a universal binary. So `packaging/macos/build.sh` and the release build
|
|
each choose one identifier for everything they build in a run. A build outside packaging gets a
|
|
new identifier when the core crate or `VERSION` changes, which is as often as cargo runs the
|
|
script again without forcing every build to relink. The daemon reports it as
|
|
`ServerInfo.build_id` (7), protocol 1.3; an older daemon leaves it empty, which matches nothing.
|
|
|
|
**One release, several packages.** The `.deb`, `.rpm` and AppImage of a release hold the same
|
|
binary and so the same identifier. A user who installs two of them has one build twice, and no
|
|
daemon to update, so neither window replaces the other's daemon. Comparing the registered path as
|
|
well was considered and left out: it would restart the daemon every time the other copy was
|
|
opened, to run the same code.
|
|
|
|
**Why reinstall rather than restart.** Restarting runs whatever the service is registered to
|
|
run, which after a second install is the other copy. Registering this copy and then starting it
|
|
is what makes the daemon the one the window came with, wherever it is.
|
|
|
|
**Asking instead of acting.** The first version replaced the daemon as soon as the window
|
|
connected. That needed guards against a loop: two windows of different builds each replacing
|
|
the other's daemon when they reconnected, and a replacement that failed being tried again on
|
|
every reconnect. The owner asked for a notice with an Update now button instead. Nothing is then
|
|
restarted unless someone asks, which removes the loop rather than guarding it, and leaves the
|
|
few seconds without MIDI to be timed by the person at the machine.
|
|
|
|
**What the notice says.** "The daemon is outdated" for an older version or one too old to say,
|
|
"a different build" for the same version, and "newer than this window" for a newer one, where the
|
|
button reads Use this version: an older program in a newer one's place may not read the
|
|
configuration it wrote, which is refused with `SchemaTooNew`. The button is offered only when the
|
|
service manager says it is running the daemon and the window is on the service's own socket.
|
|
Otherwise there is no registration that says how the daemon was started, and the notice says to
|
|
restart it from this copy.
|
|
|
|
**While the daemon is being replaced** the old one still answers for a moment. The window does
|
|
not reconnect until the replacement has finished, or it would put the notice straight back.
|
|
|
|
**Replacing under each service manager.** `ServiceManager::replace` stops the daemon, installs
|
|
the registration and starts it. Stopping comes first so the service manager stops the process it
|
|
started. launchd is the exception: installing boots the old job out, which stops its daemon, and
|
|
stopping it beforehand only has launchd start the old one again in between.
|
|
|
|
**launchd, found in use (2026-10-04).** Update now on two Macs running macOS 15.6.1 stopped the
|
|
daemon and failed with `Bootstrap failed: 5: Input/output error`, leaving the plist written and
|
|
no job loaded. `launchctl bootout` returns once the daemon has been sent SIGTERM, and the job
|
|
stays in the domain until the daemon exits; `bootstrap` is refused until then. Reproduced on
|
|
macOS 27.0.1 with a scratch job that takes 3 s to exit: `bootout` returned in 0.01 s, `print`
|
|
went on answering with `state = SIGTERMed` and `bootstrap` failed with error 5 for 3.2 s, then
|
|
`print` exited 113 and `bootstrap` succeeded. The backend now polls `print` after `bootout` until
|
|
the job has left, for up to 25 s, which covers launchd's 20 s before SIGKILL. The Start button
|
|
then ran `kickstart` on a job that was not loaded, so `start` now bootstraps the plist first when
|
|
launchd does not hold the job. The stand-in `launchctl` in `tests/service_lifecycle.rs` behaves
|
|
the same way, and the test failed against the old backend.
|
|
|
|
**The App Store build** has no service. The app starts its helper, or attaches to one already
|
|
answering. It now asks an attached daemon for its build and, when it is another, asks it to stop
|
|
and starts its own.
|
|
|
|
**Checked live** on Arch Linux under systemd, with two builds of the tree given the identifiers
|
|
`aaaaaaaa-…` and `bbbbbbbb-…`. See T253 for what was seen.
|