When someone leaves the company, their Mac comes back to IT with their working files still on it. We keep an archival copy for a retention window, then wipe the machine. The first version of that process was a technician running rsync by hand and hoping nothing was missed.
What I have now is a small pipeline of Bash scripts: a guided SwiftDialog front end for the technician, an rsync engine that recovers from hung reads, an encrypted DMG with an audit record, and a background worker that moves finished archives to a NAS and only deletes local copies after verifying them. The latest build is 1.9.1.
The Shape of It
The pipeline has three stages, each with its own script:
computerImaging.shruns on the technician’s Mac. It collects metadata, copies the source volume to a staging folder, builds the DMG, writes the audit record, and queues the archive.syncBackupsToNAS.shmounts the NAS, copies queued archives, verifies them, and cleans up.rsyncAsUser.shis the launchd entry point that runs stage two as the console user.
A job record is the only thing that connects stage one to stage two: a small JSON file in a queue directory.
jq -n \
--arg source_archive "${SUB_FOLDER}" \
--arg basename "${BASENAME}" \
--arg created_at "$(date '+%F %T')" \
'{source_archive: $source_archive, basename: $basename, created_at: $created_at}'
Early versions moved the finished archive into a watched folder to trigger the sync. With archives in the hundreds of gigabytes, that meant a second full copy on the internal disk. Now the archive stays where it was built (usually an external drive) and the queue entry just points at it. If the drive is unplugged when the worker wakes up, that job is deferred and the others proceed.
The Technician Experience
The front end is a single SwiftDialog form with five required fields: the source volume, a staging folder, the technician’s name, the serial number (entered twice to confirm), and the outgoing employee’s username. Regexes enforce the firstname.lastname format before the script runs, and the script re-validates everything afterward, since a dialog is not a trust boundary.
The staging folder gets the most scrutiny. It can’t be inside the source, the source can’t be inside it, it can’t be the NAS mount, and it has to be on a local or directly attached filesystem:
is_local_staging_path() {
local filesystem
filesystem=$(df -Pk "$1" 2>/dev/null | awk 'NR == 2 { print $1 }')
[[ -n "${filesystem}" && "${filesystem}" != //* && "${filesystem}" != smbfs:* && "${filesystem}" != nfs:* ]]
}
Every later stage (copy, app check, DMG creation) shows a progress dialog with a Cancel button. Cancel writes a private marker file, and the main loop notices it and terminates the active worker. A cancelled run keeps its partial data, because deleting a multi-hundred-gigabyte tree from a cancel button would be slow and irreversible.
rsync That Survives Bad Disks
The copy stage is where most of the work went. A Mac that has been in use for years has files that hang reads: cloud-provider placeholders, sockets, locked databases. rsync’s --timeout=60 only covers network stalls, and a local filesystem can block forever in a directory traversal.
The loop wraps rsync with a watchdog. A monitor records destination growth every 30 seconds and touches an activity marker when the backup folder gets bigger, so a long single-file copy isn’t mistaken for a stall. If neither the rsync log nor the marker changes for 15 minutes, the script kills that attempt, finds the last active path, adds it to an exclude list, and resumes. It will do this up to five times, and every skipped path is listed in the audit record.
Exit codes get classified rather than trusted:
- 24 (vanished files) is accepted.
- 23 (partial transfer) is accepted only if the error log shows source-side failures and no receiver errors, mkstemp failures, or “No space left on device.” Otherwise a destination problem could produce an archive that looks complete and silently isn’t.
- 30 (timeout) triggers a skip of the nearest common directory among the timed-out paths, so one locked database doesn’t cause a retry loop through its siblings.
- A full staging drive stops cleanly, keeps the partial copy, and tells the technician to resume on a larger drive. Rerunning the same backup resumes from the existing folder.
The excludes cover what shouldn’t be archived: /private/var/folders, Spotlight indexes, VM disk images, Docker’s raw disk file, and so on. /Applications is excluded from the main copy because those apps are reinstallable. Instead, the script runs codesign --verify against every app in /Applications and /Applications/Utilities, then copies only the ones that fail verification, using --files-from with null-delimited input. Unsigned apps are the ones you can’t get back from a vendor.
The Archive and the Audit Record
After the copy, hdiutil builds an AES-256 encrypted, compressed image. The password comes from the Keychain and goes in over stdin so it never appears in a process list:
echo -n "$DMG_PASSWORD" | hdiutil create -encryption AES-256 -format UDZO \
-stdinpass -srcfolder "${BACKUP_FOLDER}" "${DMG_NAME}" &
Next to the DMG, the script writes a plain-text audit record: technician, serial, outgoing username, start and end times, rsync exit code and summary statistics, the DMG’s SHA-256 and size, any auto-skipped paths, any unsigned apps, and every installed app with its version. The raw backup folder is deleted only after hdiutil succeeds and the image exists. Slack gets a message at each major step, and when a backup finishes the Mac announces it aloud with say, which is more effective in an IT room than a notification.
The NAS Queue
Stage two is launchd-driven. The daemon fires on a WatchPaths change to the queue directory, on a 15-minute StartInterval, and at load, with a 15-second throttle:
<key>WatchPaths</key>
<array>
<string>/Users/Shared/EndpointImagingRuntime/NASQueue</string>
</array>
<key>StartInterval</key>
<integer>900</integer>
<key>ThrottleInterval</key>
<integer>15</integer>
That path choice came from a bug. Logs and the lock originally lived in the watched directory, so every log write retriggered the worker. Runtime files now live outside the queue.
The worker takes an atomic mkdir lock, mounts the share with mount_smbfs -N (credential from the console user’s Keychain, which is why the entry script drops privileges with sudo -u), runs a write test, and syncs. Local copies are deleted only after two checks pass: an rsync dry run comparing sizes, then a SHA-256 comparison of each DMG against the checksum recorded in its audit file on the source side. It unmounts only a share that it mounted itself.
A separate script, recoverOverflowBackupToNAS.sh, handles the case where a staging drive filled up. It is deliberately opt-in, never touches the original partial backup, and checks NAS capacity first.
Packaging
Everything ships as one signed package built with munkipkg. scripts/package.sh generates build-info.json on each run, adds the Developer ID Installer identity if one is in the keychain, and optionally notarizes. The Makefile’s prepare target copies the production scripts into the payload and strips dialog_test.sh, ._* files, and extended attributes. The postinstall script boots out and re-bootstraps the daemon, then fails the install if launchctl print can’t see it, so a bad deploy is loud.
The version history in the script header is a decent record of how the work went: roughly 25 releases between September 16 and 23, most of them fixes found by running real backups. The build/ folder holds well over a hundred packages.
What I’d Do Differently
I kept an IMPROVEMENTS.md after a hardening pass, and its open items are still the right list:
- No automated tests. The scripts are checked with
bash -nand ShellCheck, but nothing exercises the retry logic. The note suggests mockingrsync,hdiutil,mount_smbfs,security, anddialog. Given how much logic sits in exit-code handling, this is the biggest gap. - Source/payload drift. The Makefile copies live scripts into a root-owned payload, and the doc notes the payload once held stale scripts. A CI preflight that runs ShellCheck, a secret scan, and a source-versus-payload comparison would catch that.
- Secrets in source. Earlier versions embedded credentials in scripts. They now come from the Keychain and are provisioned from environment variables, but anything that was ever in source or a built package needs rotating. That is a lesson for the next project: start with the Keychain.
- Retention and monitoring. Nothing yet alerts on queue items that sit longer than 24 hours, low disk space, or failed jobs. I’m also unsure whether NAS folders should be exact mirrors (
--delete-delay) or append-only. - Compression cost. UDZO saves space but can be CPU-bound. A lower zlib level may be a better trade if the NAS has room.
- Bash 3. macOS ships Bash 3.2, which already forced a sentinel element in an array to survive
set -u. A language with real data structures would make this code shorter, even if the shell-first approach made it easy to start.