Administering

Upgrading safely

An upgrade replaces the program and leaves your data where it is. This page says what Draughtsman does to make that safe, what it proves on every build, what it cannot do for you, and the four things to do before you start.

The promise, and where it stops

Upgrading does not wipe or corrupt your data, and on SQLite an upgrade that fails is undone automatically. That is backed by tests that run on every build. It is not a promise for Postgres (there is no automatic rollback, you take a pg_dump first), for a data folder you edited by hand, or for SQL Server (not supported yet). The limits are listed below.

The four things to do every time

  1. Read the release notes (the Changelog). Every release has a Breaking changes section, or says "None". A patch release (0.1.x) never needs more than the usual steps: it fixes bugs, may add settings and purely additive database migrations, and never removes, renames or narrows anything an earlier release wrote. A minor release adds features and may deprecate. A major release may remove what a deprecation left behind, and names the oldest version it can upgrade from.
  2. Back up. SQLite: draughtsman backup --out /srv/backups/; the running server need not stop. Postgres: take the pg_dump that draughtsman backup prints. See Backup, restore and upgrading. Never back up a running SQLite instance by copying data.db.
  3. Run draughtsman config check with the new program, against the same data folder and the same environment the service has.
  4. Decide where the new version goes. Keep the data folder outside the program folder (DRAUGHTSMAN_DATA, a Docker volume at /data). The program can then be at any path.

Step 1: draughtsman config check

It starts nothing and changes nothing. It reports what starting this version on this data folder would do, line by line as OK, INFO, WARN or REFUSE; only REFUSE makes it exit non-zero (1; 3 when the database is newer than this build). --json prints the same report for a script. This is a real run of the current build on a data folder an earlier build had written (a throw-away fixture from the repository's upgrade tests):

$ draughtsman config check
Draughtsman config check: this is version 0.1.0 (build 2026-10-01)
Data folder: /srv/draughtsman/data

Settings
  OK     draughtsman.yaml is read in full; every setting in it is one this version knows.

Data folder
  OK     /srv/draughtsman/data is writable by the account running this command (the server needs that to migrate, take its copy and save).

Database
  OK     /srv/draughtsman/data/data.db passes SQLite's integrity_check (9 tables, 12 documents, 3 users).
  INFO   1 of 8 migrations are pending and would be applied at start: 20261001092923_AddPasswordResetTokens.
  INFO   Before migrating, the server copies the database to data.db.pre-migration-<time> beside it, checks the copy opens and passes integrity_check, and keeps the newest 3 such copies. If a migration fails or the result loses rows, the copy is put back and the server refuses to start (exit code 4).
  OK     There is room for the copy: the database is 1 MiB and 10969 MiB is free.

Licence
  WARN   /srv/draughtsman/data/licence.key does not verify with this build (NoVerificationKey: This build carries no licence verification key, so a licence cannot be checked.). The server runs as an unlicensed evaluation, fully working. Nothing about documents depends on the licence.

Key ring
  INFO   keyRing.protection is 'none'; keys/ holds 1 key file(s): 1 plain text, 0 protected, 0 damaged.
  OK     Protection is switched off on purpose; the ring stays as plain text.

Stored secrets
  WARN   The stored AI API key cannot be read. It was written by an earlier version that tied it to the folder the program ran from, ... To recover it, start the program once from the folder it ran from before, or name that folder in keyRing.legacyInstallPaths in draughtsman.yaml (or the environment variable DRAUGHTSMAN_LEGACY_INSTALL_PATHS) and restart ... Or enter it again in Admin > AI: nothing else is affected.

Result: this version would start on this data folder (2 warning(s)).        (exit 0)

The licence warning is because this build carries no production licence key yet (see Licensing); the stored-secrets warning is the one honest limit described below, shown here for real. The same command also warns about an unknown or mistyped setting (server.prot: did you mean server.port?), a key ring whose protection is not available in your shell, and a file that is not valid YAML (the one thing that does stop the server).

Step 2: replace the program, not the data

  • Release archive or service: stop the service, unpack the new archive over (or in place of) the old program folder, leave the data folder alone, start the service. If you kept the archive's default data/ folder beside the executable, unpack over the old folder (the archive contains no data/) or move data/ out first. On Windows the new folder is named for the new version, so run service install again from it.
  • Docker: build or load the new image under a new tag, stop and remove the old container, run the new tag against the same volume.

Step 3: start it, and what it does on SQLite

Before it changes anything the server copies data.db to data.db.pre-migration-<time> beside it (after folding in the write-ahead log) and proves the copy: the same size, it opens, it passes SQLite's integrity_check. Then it applies each pending migration in a transaction, and then checks the result: integrity_check must be ok and no table may hold fewer rows than in the copy. Only then does it prune older copies (it keeps the newest three, retention.preMigrationCopies). Real log lines from the run above:

info: draughtsman[0]
      Copied SQLite database to /srv/draughtsman/data/data.db.pre-migration-20261001160426 before migrating; the copy opens and passes integrity_check (9 tables)
info: draughtsman[0]
      Applying 1 storage migration(s) to sqlite: 20261001092923_AddPasswordResetTokens
info: draughtsman[0]
      Post-migration check passed: SQLite integrity_check is ok and no table has fewer rows than in /srv/draughtsman/data/data.db.pre-migration-20261001160426

Settings in draughtsman.yaml are never rewritten by an upgrade; a setting the new version adds takes its default. An old file keeps working: a setting a later release renames is read under its new name, one it removes is ignored, and each is a warning that names what to do, never a failure. Compare your file with draughtsman.example.yaml in the new archive.

When an upgrade fails: automatic rollback (SQLite)

If a migration throws, or the post-migration checks fail, the server puts the copy back over the database (deleting the -wal and -shm files first so the failed run's log cannot be replayed), verifies the restored file is byte for byte the copy, refuses to start and exits with code 4. If the copy cannot be made (disk full, folder not writable, or the live file already fails integrity_check) the upgrade is refused before anything changes, with exit code 5. If you must upgrade anyway, take your own backup and set storage.allowMigrationWithoutBackup: true; the server then logs MIGRATING WITHOUT A BACKUP and goes ahead with no automatic way back. Switch it off afterwards.

No shipped migration has ever lost data, so the rollback cannot be shown by a real upgrade. We forced it: on a copy of the fixture we added a SQLite trigger that deletes every document the moment the pending migration is recorded, then started the server. This is the real output, and the database file's checksum was identical before and after:

crit: draughtsman[0]
      Storage migration undone: /srv/draughtsman/data/data.db was restored from /srv/draughtsman/data/data.db.pre-migration-20261001160504 and is exactly as it was before this start. Reasons: table document_versions had 9 rows before the migration and has 0 now; table documents had 12 rows before the migration and has 0 now
crit: draughtsman[0]
      The storage upgrade of /srv/draughtsman/data/data.db did not pass its checks: ... The pre-upgrade copy ... was put back automatically, so data.db is exactly as it was before this start (the copy is kept beside it) and nothing was served. Start the previous version again, and report this failure with the log.
$ echo $?
4
$ md5 data.db   # before the start, and after it
cafee9eefaf127b0de70d8bef00d5e13   cafee9eefaf127b0de70d8bef00d5e13

Postgres: the safety net is yours

The server cannot snapshot your database, so it does not copy it, and there is no automatic rollback. When an existing Postgres database has migrations to apply, the server logs a loud warning first, naming the migrations and the pg_dump to run, and goes ahead; draughtsman config check shows the same block. To go back, stop the new version, empty the schema and restore the dump into it (DROP SCHEMA public CASCADE; CREATE SCHEMA public; then pg_restore --no-owner). Do not use pg_restore --clean: it drops only what the dump holds, so a table the upgrade added blocks it and leaves a half-restored database. That was found by trying it. The migrations themselves were run against populated Postgres databases made by seven earlier builds, and a Postgres database was upgraded and rolled back from a dump on a real build; we did not repeat the Postgres runs for this site.

An older program refuses a newer database

If you start an older version on a database a newer version has migrated, it logs a critical message naming the migrations it does not know, changes nothing, serves nothing and exits with code 3. mcp --stdio, reset-admin and restore refuse the same way, and restore refuses a bundle from a newer version. We reproduced it by adding a migration row the build does not know to a copy of the fixture:

$ draughtsman config check
Database
  OK     /srv/draughtsman/fut/data.db passes SQLite's integrity_check (9 tables, 12 documents, 3 users).
  REFUSE This database was migrated by a NEWER version of Draughtsman (it has 1 migration(s) this build does not know: 29991231000000_FromTheFuture). This version would refuse to open it and exit with code 3. ...
                                                                          (exit 3)
$ draughtsman
crit: draughtsman[0]
      This sqlite database was migrated by a newer version of Draughtsman than this one (it has 1 migration(s) this build does not know: 29991231000000_FromTheFuture). ... so it was not opened and nothing was changed. Either start the newer Draughtsman version again, or downgrade by restoring what was saved before the upgrade. ...
                                                                          (exit 3)

Builds before the commit that introduced this refusal (257d224) do not have it and would try to run on the newer schema. Never go back across it without restoring the pre-upgrade copy or a backup first. (Overpass also ran the real older archive and the older Docker image over a database the new build had migrated, and both refused with exit code 3; we did not rerun that.)

Going back

  1. Stop the new version.
  2. SQLite: delete data.db-wal and data.db-shm first, then copy the newest data.db.pre-migration-* over data.db. A write-ahead log left behind by the migrated database would otherwise be replayed onto the restored file and silently return the instance to the new schema. Or draughtsman restore <bundle> --overwrite, which sets the old data aside in pre-restore-<time>/.
  3. Put the older program back, at the same path as before, and start it.
  4. If the new version re-protected the stored AI key or SMTP password (the log said so, and backups/secrets-<time>/ exists), copy ai/settings.json and email/settings.json from that folder over the ones in the data folder first. An older program only reads them from the folder it ran from, and otherwise shows them as needing to be entered again. Nothing is lost either way.

Anything saved after the upgrade is lost in a downgrade; there is no way to run an older version over newer data.

Stored secrets, and moving the program

The stored AI API key and the saved SMTP password are encrypted under a key ring in the data folder. ASP.NET Core's Data Protection also binds what it encrypts to an "application name". Older builds left that at the framework default, which is the folder the program ran from, so moving the program or unpacking a new version into a new folder made both secrets unreadable and an administrator had to type them again. The name is now the fixed word Draughtsman, so a secret depends on the key ring in the data folder alone: any path, a new folder beside the old one, a Docker image whose program is always /app. This is tested on the data folder of every earlier build, and on release archives and a Docker volume, including a second move to a third folder with no fallback at all.

The one honest limit: installs from before this fix

A data folder written by an earlier build has its secrets tied to the folder that build ran from, and that folder is part of the encryption: it cannot be recovered from the data, only guessed or told. The first start of this version tries the folder it runs from, the folders it has recorded, any you name, the folder the data folder sits in, the folders beside it, old log files and the recommended install locations, reads the secret and re-protects it (keeping the earlier settings file whole in backups/secrets-<time>/). If you upgrade in place, nothing special is needed. If you unpack elsewhere, run config check first: its Stored secrets section says per secret whether it can be read. If every guess fails (the old folder is gone and the program moved at the same time), nothing is changed or dropped, Admin > AI, Admin > Email and config check say which secret and why, and naming the old folder once fixes it:

keyRing:
  legacyInstallPaths:
    - /opt/draughtsman-0.1.0     # the folder the OLD program ran from

or set DRAUGHTSMAN_LEGACY_INSTALL_PATHS. Failing that, enter the key again under Admin > AI; nothing else is affected. Our config check run above showed this warning for real: the fixture was written by an old build at a different path. 0.1.0 has not been cut as a numbered release yet, so this concerns copies run from earlier builds.

Also give the new version the same key-ring protection: with keyRing.protection: key, keep DRAUGHTSMAN_KEY_ENCRYPTION_KEY (or the key file) in the service's environment; with a certificate, the certificate and its password variable. The server still starts without it, but the stored AI key is unreadable until it is back. Browser sessions carry over. A tab signed in before the first start of this version holds a CSRF token the new server cannot read once; its first refused save carries a fresh one and the editor retries, so no edit is lost.

What is proved, and what is not

Proved on every build (in the repository's tests)

  • The repository holds the data folder of seven earlier builds, each made by running that build and filling it through its own API: users and passwords, API tokens (plain, expiring, read-only, revoked), documents of every type including a 2,000-node diagram and a fly-through, versions, a bin, branding, organisation themes, AI settings with a stored key and SMTP password, a licence file and security events.
  • Every test run starts the current server on a copy of each, on SQLite and on Postgres, and checks every document is byte for byte what the old build served, every user signs in with the same password, every token still works with the same scope and expiry, and no table has fewer rows. The list of differences an upgrade is allowed to make is empty.
  • No migration can silently destroy data. A test scans every migration of both providers and fails on a dropped table or column, a rename, a narrowing or type-changing column, a NOT NULL column with no default, or raw SQL that deletes, truncates or drops, unless a written justification and the fixture that proves it are on file. The rule is expand, then contract: a release only adds, and what an older release no longer reads is dropped by a later release.
  • A release is cut only after its version has an upgrade fixture; the release script refuses without one.

Run for real on release binaries by Overpass (not repeated here)

  • An archive swap in place from an older build, with the pre-upgrade copy byte for byte the old database, and the old program then refusing the migrated database; the same sequence on a Docker volume; the stored-secrets move on release archives and on Docker; a Postgres upgrade and rollback from a dump.

Not covered, so you can decide what to do about it

  • Postgres has no automatic rollback. The migrations are tested against populated Postgres databases; the safety net around them is your pg_dump.
  • The upgrade and rollback steps on a Windows service and a Linux systemd service were not run on those service managers; the logic is tested on every platform.
  • A data folder edited by hand, files the server owns that you changed (data.db, ai/settings.json, the key ring), a folder written before the first pushed build, and SQL Server (not supported).
  • Databases that already fail SQLite's integrity_check. The upgrade refuses to touch them; restore a backup first.
  • An organisation theme's updatedAt is its file's modification time, which a copy of the folder does not keep.

Note on this page's own checks: the commands above were run on the current main build started from the repository, against copies of the repository's upgrade fixture, because a full release archive could not be built on the machine at the time (it was out of memory). The release-archive and Docker runs are Overpass's own, as the repository's upgrade guide records.