Skip to content

Releasing

The step-by-step release runbook. For the why and the shape of the branch flow, see Branching & releases; this page is the exact steps.

Run pnpm run cut-release -- X.Y.Z. The rest is automatic. You can also cut a release from GitHub Actions by running Cut Release on the develop branch with the same X.Y.Z version.

Normal release flow

  1. Make sure the intended changes are merged to develop.
  2. From a clean checkout at the latest origin/develop, run:
git checkout develop
git pull --ff-only origin develop
pnpm run cut-release -- X.Y.Z
  1. Review the generated release/X.Y.Z -> main PR.
  2. Merge the release PR to main with a merge commit. This starts Deploy - Production, which ships every component in one run and cuts the GitHub release when it finishes.
  3. Let the Back Merge Main to Develop workflow open or update sync/main-to-develop.
  4. Merge the sync PR to develop with a merge commit.
  5. Watch the deploy land. The run posts to #product-team either way; see Deployments if it fails.

Releases are cut from develop, never from main, so the release branch is grounded in the same state that was tested on staging. The post-release sync must be a real merge commit so release commits propagate back to develop and the branches do not silently diverge. Incident context: PR #976.

GitHub Actions entry point

Use Actions -> Cut Release -> Run workflow on develop. Inputs:

  • version: MAJOR.MINOR.0 without a build suffix, for example 1.50.0. See Versioning.
  • dry_run: validates and prints the planned branch, version bump and PR without pushing anything.

The workflow wraps scripts/cut-release.ts, configures the GitHub Actions bot identity, and creates the same release PR as the local script. It then runs the Jira fix-version restamp (Next -> X.Y.Z) itself: a branch pushed with the workflow's own token never delivers the create event that jira-restamp-fix-version.yml listens for. A local pnpm run cut-release pushes as you, so the create trigger covers it.

Once the branch and PR exist, the workflow posts the version, branch and PR link to #tech in Slack (secret SLACK_WEBHOOK_TECH, a webhook bound to that channel). A local pnpm run cut-release does not post, so prefer the workflow, or announce the cut by hand.

Versioning

Release versions follow semver: MAJOR.MINOR.PATCH.

Part Bumped for Example
MAJOR A new platform (for example Platform 2), or the same platform with breaking changes. 1.49.0 -> 2.0.0
MINOR A normal release with new features. Reset PATCH to 0. 1.49.0 -> 1.50.0
PATCH A hotfix to what is on main. 1.49.0 -> 1.49.1

cut-release only accepts MAJOR.MINOR.0. A hotfix bumps PATCH in both pubspecs on its hotfix/* branch; a build-number-only bump no longer passes the release guard.

Until October 2026 releases were numbered 1.1.N, with hotfixes as build numbers (1.1.49+2). Under the new scheme release 1.1.49 is 1.49.0, its hotfix is 1.49.1, and the next release is 1.50.0. The guard compares semver numerically, so 1.50.0 is greater than 1.1.49 and the move needs no special case.

The version of record lives in both Flutter app pubspecs, which the release script writes as <version>+1:

  • apps/perci-platform-members/pubspec.yaml
  • apps/perci-platform-clinicians/pubspec.yaml

The +1 build number is intentionally stable for web releases. The release version must be strictly greater than the version currently on main; the release guard rejects downgrades and stale version bumps.

What a production deploy publishes

A successful Deploy - Production run produces one GitHub release covering the whole train (backend functions, Firestore rules, Medplum bots, the Web Professionals portal and both apps), replacing the three per-component releases the old per-app workflows cut. The body is generated by scripts/deploy/release-notes.ts and the tag is derived from the members pubspec version.

Re-running a production deploy does not fail on an already-published tag: an existing release is left untouched. The run also closes a single GitHub deployment record for the environment. See Deployments.

Flutter web source maps

Both apps are built with --source-maps, and each deploy job uploads main.dart.js.map to Datadog with datadog-ci flutter-symbols upload so prod errors arrive as Dart file names and line numbers instead of main.dart.js offsets.

Three invariants keep that working:

  • The upload --version must equal the version the RUM SDK reports at runtime. The SDK sends the pubspec version, the build number and the commit the bundle was built from, joined by - (datadogReleaseVersion in datadog_reporting.dart; the commit arrives as the BUILD_ID dart-define). The workflow builds the same string: the pubspec version with + replaced by -, then -<commit SHA>. A mismatch is silent: Datadog accepts the upload and never applies it. --service-name must likewise match the SDK's service. The commit is there so that every build has its own version, which the next invariant depends on.
  • The first map uploaded for a version is the one Datadog keeps. Re-uploading under the same service and version does not replace it, so a map cannot be corrected after the fact: whatever has to be in it (see the embedding below) must be there before the upload step, and re-running a deploy of the same commit leaves the original map in place.
  • The maps are never served publicly. firebase.json ignores **/*.map in every hosting target, so the .map files exist only in the build artefact and in Datadog. They stay in build/web for the upload step, which runs after the hosting deploy in the same job.

datadog-ci exits 0 even when it skips every map, so each upload step asserts that main.dart.js.map was uploaded and raises a workflow error if it was not. The step is continue-on-error, so a Datadog outage annotates the run without failing a shipped release. The two skipped maps it always reports (flutter.js.map, hls.min.js.map) are sourceMappingURL references to files Flutter does not emit, and are expected.

dart2js writes source paths but no sourcesContent, and Datadog renders code snippets in stack traces only from sourcesContent. Each build job therefore runs scripts/frontend/embed-sourcemap-sources.mjs right after flutter build web, while the repo, pub cache and Flutter SDK are still on the runner. It grows main.dart.js.map from about 13 MB to about 70-75 MB (80-87 MB once the padding below is added, which is the file that is uploaded), well inside Datadog's 500 MB limit (map plus minified file), and costs users nothing since the map is never served. It is continue-on-error too: if a source cannot be found it still writes the rest, raises a workflow error, and the upload goes ahead with whatever was embedded.

The same script pads the mappings. dart2js records few positions per generated line, and Datadog leaves a frame as a raw main.dart.js offset when the nearest recorded position is more than a few columns before it. The script repeats each position every 4 columns up to the next one, which changes no lookup (a consumer already answers with the nearest earlier position) and adds about 10 MB to the map.

One kind of frame is fixed in the app instead, because the map cannot help with it. The SDK forwards the browser's stack string as is, and Datadog only maps frames written at name (url:line:column). A frame for an anonymous function is at url:line:column, and dart2js compiles every async body to one, so the throw site of every asynchronous error stayed raw. DatadogErrorReporter.withNamedFrames rewrites those frames before reporting; errors sent straight to DatadogSdk.instance.rum skip it.

QA fixes

During QA, fixes land on develop first. Forward-port the same fix to the open release/X.Y.Z branch so the release contains the reviewed change and develop stays the source of truth.

Hotfixes

Use a hotfix when a fix must reach production before the next normal release. The rules for main PRs (branch prefix, version bump) are in Branching & releases.

A hotfix is not the fastest way out of a bad deploy

A hotfix still has to go through review, CI and a full production pipeline run. When production is actively broken, roll back first and fix forward after. See the Rollback runbook.

For urgent production incidents:

  1. Branch from main as hotfix/PPL-NNNN.
  2. Keep the fix minimal.
  3. Open a PR to main.
  4. Merge with a merge commit after review and required checks pass.
  5. Let the automatic sync/main-to-develop PR carry the hotfix back to develop.

For non-urgent hotfixes, prefer fixing develop first:

  1. Branch from develop and merge the fix to develop through the normal PR path.
  2. Branch from main as hotfix/PPL-NNNN.
  3. Cherry-pick the minimal fix commit from develop onto the hotfix branch.
  4. Open the hotfix PR to main and merge it with a merge commit.
  5. Let the automatic sync/main-to-develop PR reconcile the production merge back to develop.

Before cherry-picking from develop, check the fix does not depend on unreleased develop-only changes. If it does, make a smaller hotfix directly from main instead.

Required repository settings

main keeps its protections and requires:

  • Pull requests; no direct pushes.
  • Required status check: Release PR Guard.
  • Merge commits allowed for release/*, hotfix/* and sync PRs.
  • Squash remains the default for ordinary PRs to develop.
  • PRs to main must come from release/*, hotfix/* or sync/main-to-develop*. The versioned enforcement lives in scripts/release-pr-guard.ts; if GitHub branch rulesets are available, mirror the same source-branch allow-list in the main ruleset.

DORA metrics

Four numbers tell us whether our delivery is getting better: deployment frequency, lead time for change, change failure rate and time to restore. All four are derived by Datadog from two event streams that the deploy pipeline emits.

What CI emits

The finish job of .github/workflows/deploy.yml calls the .github/actions/dora-event composite action once per release, only on a successful production run:

Field Value reported
DORA service perci-platform
Version members app version from pubspec.yaml, + rewritten to -
git.commit_sha the commit the release deployed

One event, because one run is one release. A production run ships the whole platform in one train, so there is one thing to count. Until PPL-3419 this job emitted three events, one each for perci-platform-backend, members-app-frontend and clinicians-app-frontend. That modelled the per-app deploy workflows we no longer have: all three carried the same commit, the same started_at and the same outcome, so they were the same event three times and deployment frequency read three times the real release rate.

The version is reported as 1.1.28-1, not the 1.1.28+1 the pubspec carries: the build jobs normalise the separator because + is not safe everywhere the version is passed through. Expect the hyphen form when correlating a DORA event against tooling that reports the raw pubspec value. It comes from the members pubspec because the GitHub release tag is derived from the same file, so one DORA event maps to exactly one GitHub release. The release guard requires both app pubspecs to be bumped together, so the clinicians version is always the same string. The commit is reported separately as git.commit_sha, which is what lead time for change is computed from, so nothing is lost by not putting the SHA in version.

The trade-off. perci-platform is not an APM or RUM service, so Datadog cannot overlay these deployments onto a service dashboard the way it could when the DORA service names matched. That is deliberate: a correct headline number beats a correlation we rarely used. Version tagging and source-map uploads are a separate mechanism and still report per service.

Events are tagged with the run's environment. Only production runs emit, so in practice every event is env:prod.

The action sends two kinds of event to the EU Datadog site (api.datadoghq.eu):

  • Deployment event (POST /api/v2/dora/deployment) on every successful deploy. This feeds deployment frequency and lead time for change. The started_at timestamp is captured by the Record deploy start step, the first step of the begin job, before any build or deploy work in the run. Capture it later and both the deployment duration and the lead time come out wrong.
  • Change-failure event (POST /api/v2/dora/incident) when the deploy is a hotfix. This feeds change failure rate and time to restore. Because production deploys are triggered by a push to main, the ref is always main, so a hotfix is recognised from the merge commit GitHub writes for a hotfix/* PR (Merge pull request #N from <org>/hotfix/...). That is reliable because PRs into main may only come from release/* or hotfix/* and always land as a merge commit, never a squash. See Branching & releases.

The failure window is reported as previous production commit to hotfix deployed: the incident started_at is the commit time of HEAD^ (the version that was live and broken) and finished_at is the moment the hotfix finished deploying. That is an approximation, because it assumes the breakage arrived with the previous production deploy, and it is the best signal CI has today. If HEAD^ cannot be resolved (a shallower checkout than fetch-depth: 2) the start falls back to the deploy start rather than to finished_at, so a failed lookup cannot report a zero-length incident and flatter a time-to-restore figure.

Robustness

Emission is best-effort and can never break a release:

  • .github/actions/dora-event/dora-event.sh never exits non-zero. A missing key, a Datadog outage, a 4xx or a malformed input logs a GitHub warning annotation and returns 0.
  • The calling step is additionally guarded with continue-on-error: true.
  • The steps are gated on the pipeline outcome the finish job resolves from every other job, so nothing is emitted for a run that failed or was cancelled anywhere along the way.
  • Missing DATADOG_API_KEY skips emission with a warning, matching how the sourcemap upload steps behave.

No new repository secret is needed: the action authenticates with the existing DATADOG_API_KEY, the same secret the sourcemap and Cloud Run instrumentation steps already use. The DORA intake authenticates with the API key alone, so no application key is sent.

Still to do by hand

The event pipeline is only half the deliverable. These parts live in the Datadog console and on Confluence and are not in this repo:

  1. Keep every service's DORA source set to API. Datadog can also derive deployments from APM version tags, and a service with both sources on counts each release twice. That is what PPL-3419 found: clinicians-app-frontend had APM deployment tracking on, so it produced a duplicate source:apm_deployments event for every release, plus a third for the web bundle, whose RUM version carries the BUILD_ID suffix and reads as a separate deployment. Check this whenever a DORA number looks inflated: filter the deployments list by source and every event should say api.
  2. Confirm events are arriving. After the next production deploy, open Datadog → Software Delivery → DORA Metrics and check exactly one deployment event exists for perci-platform. Confirm the exact metric and event field names in the console before wiring widgets to them.
  3. Build the dashboard. One dashboard, four widgets (deployment frequency, lead time for change, change failure rate, time to restore) over service:perci-platform. If Datadog offers an out-of-the-box DORA dashboard, clone that rather than hand-building it.
  4. Record the baseline. Once a full four weeks of events have accumulated, write the starting value of each of the four metrics onto the engineering strategy page in Confluence, so later improvement is measured against something rather than asserted.
  5. Review on a cadence. Recommended: review the four numbers monthly at the engineering review, and re-baseline quarterly. Treat a metric that has not moved for two quarters as a signal that the underlying delivery problem was misidentified, not that the metric is wrong.

Known gaps

  • Hotfixes are the only change-failure signal today. A production incident fixed by a normal release, or one resolved without a code change, does not produce an incident event, so change failure rate and time to restore are floors, not exact figures.
  • Canary auto-rollback is not wired. That signal belongs to PPL-3149 and does not exist yet. The seam is already in the action: call it with emit_incident: 'true' plus incident_name and incident_started_at (the rollback knows when the bad version went out, so it can report a far more accurate failure window than the HEAD^ approximation above).
  • The pipeline's own rollback-backend job does not emit an incident event either. A verification failure that triggers an automatic Cloud Run rollback is invisible to change failure rate today. It uses the same seam as the item above.
  • Staging deploys emit nothing. All emission is gated on inputs.environment == 'production', so there is no staging series to filter out. An env:staging deployment event in Datadog did not come from this pipeline, and is a sign that a service has an APM-derived DORA source switched on.
  • There is no per-service breakdown any more. Deployment frequency, lead time, change failure rate and time to restore all describe the release train, not the backend or either app individually. Nothing is lost today, because the components never ship apart. If they ever do, this is the first thing to revisit.

Troubleshooting

release must be cut from the exact origin/develop commit

git checkout develop
git pull --ff-only origin develop
pnpm run cut-release -- X.Y.Z

version must be greater

Pick the next semver above the version currently on both app pubspecs. Do not reuse a version that has already shipped.

remote branch already exists: origin/release/X.Y.Z

An earlier cut of the same version left its branch behind, usually because the release PR was closed instead of merged. Delete the stale branch and cut again:

git push origin --delete release/X.Y.Z
pnpm run cut-release -- X.Y.Z

If the release PR is still open, do not recut. Merge develop into the release branch instead (see QA fixes).

PR branch appears to be grounded in main

Close the bad release PR and re-run pnpm run cut-release -- X.Y.Z from the latest origin/develop.

.gitignore files were added

Remove generated build output from the release branch. Common examples are Flutter Linux generated files under apps/*/linux/**, .dart_tool, and build/.

mcp/*/package-lock.json

Remove the npm lockfile. MCP packages use pnpm, and the root pnpm-lock.yaml is the lockfile source of truth.

Back-merge PR has conflicts

Check out sync/main-to-develop, merge origin/main, resolve conflicts, then push the branch. Keep the PR as a merge commit into develop.