ToolDrift: 7 failures after refresh success
ToolDrift logged seven failed deploys in 24 hours after every refresh succeeded
Seven failure reports told us ToolDrift’s nightly refresh was broken, even though every refresh had completed successfully within the same 24 hours.
The incident
ToolDrift refreshes its data on a recurring run. The run executes its tests, writes the refreshed data, and then attempts to publish the resulting pages. The publishing step can refuse to proceed when the working tree contains uncommitted source changes. That refusal protects unfinished work from reaching production.
The refusal was correct. Our interpretation of it was wrong.
The publishing script returned a nonzero exit code after refusing the dirty tree. ToolDrift’s wrapper treated that code as proof that publishing had failed. It recorded the whole run as failed, even though the tests had passed and the refresh had written its data.
What launchd received
launchd did not receive the meaning of the refusal. Its plist supplied a working directory, explicit environment variables, log destinations, and a list of program arguments. It then received only the wrapper’s final process status.
That boundary matters because an exit code carries very little context. The same nonzero value covered both a deliberate refusal before any build began and an actual build or upload failure. The wrapper collapsed those different outcomes into one failure state and handed that state back to launchd.
The wrapper around the job also counts repeated failures. Enough matching failures cause it to remove the job from launchd, which prevents the next refresh from starting. A healthy protection inside the publishing step could