Releasing¶
The step-by-step release runbook. For the why and the shape of the branch flow, see Branching & releases; this page is the exact steps.
Run pnpm run cut-release -- X.Y.Z. The rest is automatic. You can also cut a release from
GitHub Actions by running Cut Release on the develop branch with the same X.Y.Z
version.
Normal release flow¶
- Make sure the intended changes are merged to
develop. - From a clean checkout at the latest
origin/develop, run:
- Review the generated
release/X.Y.Z->mainPR. - Merge the release PR to
mainwith a merge commit. This starts Deploy - Production, which ships every component in one run and cuts the GitHub release when it finishes. - Let the Back Merge Main to Develop workflow open or update
sync/main-to-develop. - Merge the sync PR to
developwith a merge commit. - Watch the deploy land. The run posts to
#product-teameither way; see Deployments if it fails.
Releases are cut from develop, never from main, so the release branch is grounded in the
same state that was tested on staging. The post-release sync must be a real merge commit so
release commits propagate back to develop and the branches do not silently diverge.
Incident context: PR #976.
GitHub Actions entry point¶
Use Actions -> Cut Release -> Run workflow on develop. Inputs:
version:MAJOR.MINOR.0without a build suffix, for example1.50.0. See Versioning.dry_run: validates and prints the planned branch, version bump and PR without pushing anything.
The workflow wraps scripts/cut-release.ts, configures the GitHub Actions bot identity, and
creates the same release PR as the local script. It then runs the Jira fix-version restamp
(Next -> X.Y.Z) itself: a branch pushed with the workflow's own token never delivers the
create event that jira-restamp-fix-version.yml listens for. A local pnpm run cut-release
pushes as you, so the create trigger covers it.
Once the branch and PR exist, the workflow posts the version, branch and PR link to #tech in
Slack (secret SLACK_WEBHOOK_TECH, a webhook bound to that channel). A local
pnpm run cut-release does not post, so prefer the workflow, or announce the cut by hand.
Versioning¶
Release versions follow semver: MAJOR.MINOR.PATCH.
| Part | Bumped for | Example |
|---|---|---|
MAJOR |
A new platform (for example Platform 2), or the same platform with breaking changes. | 1.49.0 -> 2.0.0 |
MINOR |
A normal release with new features. Reset PATCH to 0. |
1.49.0 -> 1.50.0 |
PATCH |
A hotfix to what is on main. |
1.49.0 -> 1.49.1 |
cut-release only accepts MAJOR.MINOR.0. A hotfix bumps PATCH in both pubspecs on its
hotfix/* branch; a build-number-only bump no longer passes the release guard.
Until October 2026 releases were numbered 1.1.N, with hotfixes as build numbers
(1.1.49+2). Under the new scheme release 1.1.49 is 1.49.0, its hotfix is 1.49.1, and the
next release is 1.50.0. The guard compares semver numerically, so 1.50.0 is greater than
1.1.49 and the move needs no special case.
The version of record lives in both Flutter app pubspecs, which the release script writes as
<version>+1:
apps/perci-platform-members/pubspec.yamlapps/perci-platform-clinicians/pubspec.yaml
The +1 build number is intentionally stable for web releases. The release version must be
strictly greater than the version currently on main; the release guard rejects downgrades
and stale version bumps.
What a production deploy publishes¶
A successful Deploy - Production run produces one GitHub release covering the whole
train (backend functions, Firestore rules, Medplum bots, the Web Professionals portal and
both apps), replacing the three per-component releases the old per-app workflows cut. The
body is generated by scripts/deploy/release-notes.ts and the tag is derived from the
members pubspec version.
Re-running a production deploy does not fail on an already-published tag: an existing release is left untouched. The run also closes a single GitHub deployment record for the environment. See Deployments.
Flutter web source maps¶
Both apps are built with --source-maps, and each deploy job uploads main.dart.js.map to
Datadog with datadog-ci flutter-symbols upload so prod errors arrive as Dart file names and
line numbers instead of main.dart.js offsets.
Three invariants keep that working:
- The upload
--versionmust equal the version the RUM SDK reports at runtime. The SDK sends the pubspec version, the build number and the commit the bundle was built from, joined by-(datadogReleaseVersionindatadog_reporting.dart; the commit arrives as theBUILD_IDdart-define). The workflow builds the same string: the pubspec version with+replaced by-, then-<commit SHA>. A mismatch is silent: Datadog accepts the upload and never applies it.--service-namemust likewise match the SDK'sservice. The commit is there so that every build has its own version, which the next invariant depends on. - The first map uploaded for a version is the one Datadog keeps. Re-uploading under the same service and version does not replace it, so a map cannot be corrected after the fact: whatever has to be in it (see the embedding below) must be there before the upload step, and re-running a deploy of the same commit leaves the original map in place.
- The maps are never served publicly.
firebase.jsonignores**/*.mapin every hosting target, so the.mapfiles exist only in the build artefact and in Datadog. They stay inbuild/webfor the upload step, which runs after the hosting deploy in the same job.
datadog-ci exits 0 even when it skips every map, so each upload step asserts that
main.dart.js.map was uploaded and raises a workflow error if it was not. The step is
continue-on-error, so a Datadog outage annotates the run without failing a shipped release.
The two skipped maps it always reports (flutter.js.map, hls.min.js.map) are
sourceMappingURL references to files Flutter does not emit, and are expected.
dart2js writes source paths but no sourcesContent, and Datadog renders code snippets in
stack traces only from sourcesContent. Each build job therefore runs
scripts/frontend/embed-sourcemap-sources.mjs right after flutter build web, while the
repo, pub cache and Flutter SDK are still on the runner. It grows main.dart.js.map from
about 13 MB to about 70-75 MB (80-87 MB once the padding below is added, which is the
file that is uploaded), well inside Datadog's 500 MB limit (map plus minified file),
and costs users nothing since the map is never served. It is continue-on-error too: if a
source cannot be found it still writes the rest, raises a workflow error, and the upload goes
ahead with whatever was embedded.
The same script pads the mappings. dart2js records few positions per generated line, and
Datadog leaves a frame as a raw main.dart.js offset when the nearest recorded position is
more than a few columns before it. The script repeats each position every 4 columns up to the
next one, which changes no lookup (a consumer already answers with the nearest earlier
position) and adds about 10 MB to the map.
One kind of frame is fixed in the app instead, because the map cannot help with it. The SDK
forwards the browser's stack string as is, and Datadog only maps frames written
at name (url:line:column). A frame for an anonymous function is at url:line:column, and
dart2js compiles every async body to one, so the throw site of every asynchronous error
stayed raw. DatadogErrorReporter.withNamedFrames rewrites those frames before reporting;
errors sent straight to DatadogSdk.instance.rum skip it.
QA fixes¶
During QA, fixes land on develop first. Forward-port the same fix to the open
release/X.Y.Z branch so the release contains the reviewed change and develop stays the
source of truth.
Hotfixes¶
Use a hotfix when a fix must reach production before the next normal release. The rules for
main PRs (branch prefix, version bump) are in
Branching & releases.
A hotfix is not the fastest way out of a bad deploy
A hotfix still has to go through review, CI and a full production pipeline run. When production is actively broken, roll back first and fix forward after. See the Rollback runbook.
For urgent production incidents:
- Branch from
mainashotfix/PPL-NNNN. - Keep the fix minimal.
- Open a PR to
main. - Merge with a merge commit after review and required checks pass.
- Let the automatic
sync/main-to-developPR carry the hotfix back todevelop.
For non-urgent hotfixes, prefer fixing develop first:
- Branch from
developand merge the fix todevelopthrough the normal PR path. - Branch from
mainashotfix/PPL-NNNN. - Cherry-pick the minimal fix commit from
developonto the hotfix branch. - Open the hotfix PR to
mainand merge it with a merge commit. - Let the automatic
sync/main-to-developPR reconcile the production merge back todevelop.
Before cherry-picking from develop, check the fix does not depend on unreleased
develop-only changes. If it does, make a smaller hotfix directly from main instead.
Required repository settings¶
main keeps its protections and requires:
- Pull requests; no direct pushes.
- Required status check: Release PR Guard.
- Merge commits allowed for
release/*,hotfix/*and sync PRs. - Squash remains the default for ordinary PRs to
develop. - PRs to
mainmust come fromrelease/*,hotfix/*orsync/main-to-develop*. The versioned enforcement lives inscripts/release-pr-guard.ts; if GitHub branch rulesets are available, mirror the same source-branch allow-list in themainruleset.
DORA metrics¶
Four numbers tell us whether our delivery is getting better: deployment frequency, lead time for change, change failure rate and time to restore. All four are derived by Datadog from two event streams that the deploy pipeline emits.
What CI emits¶
The finish job of .github/workflows/deploy.yml calls the .github/actions/dora-event
composite action once per release, only on a successful production run:
| Field | Value reported |
|---|---|
| DORA service | perci-platform |
| Version | members app version from pubspec.yaml, + rewritten to - |
git.commit_sha |
the commit the release deployed |
One event, because one run is one release. A production run ships the whole platform in one
train, so there is one thing to count. Until PPL-3419 this job emitted three events, one each
for perci-platform-backend, members-app-frontend and clinicians-app-frontend. That
modelled the per-app deploy workflows we no longer have: all three carried the same commit, the
same started_at and the same outcome, so they were the same event three times and deployment
frequency read three times the real release rate.
The version is reported as 1.1.28-1, not the 1.1.28+1 the pubspec carries: the build jobs
normalise the separator because + is not safe everywhere the version is passed through.
Expect the hyphen form when correlating a DORA event against tooling that reports the raw
pubspec value. It comes from the members pubspec because the GitHub release tag is derived from
the same file, so one DORA event maps to exactly one GitHub release. The release guard requires
both app pubspecs to be bumped together, so the clinicians version is always the same string.
The commit is reported separately as git.commit_sha, which is what lead time for change is
computed from, so nothing is lost by not putting the SHA in version.
The trade-off. perci-platform is not an APM or RUM service, so Datadog cannot overlay
these deployments onto a service dashboard the way it could when the DORA service names matched.
That is deliberate: a correct headline number beats a correlation we rarely used. Version
tagging and source-map uploads are a separate mechanism and still report per service.
Events are tagged with the run's environment. Only production runs emit, so in practice every
event is env:prod.
The action sends two kinds of event to the EU Datadog site (api.datadoghq.eu):
- Deployment event (
POST /api/v2/dora/deployment) on every successful deploy. This feeds deployment frequency and lead time for change. Thestarted_attimestamp is captured by theRecord deploy startstep, the first step of thebeginjob, before any build or deploy work in the run. Capture it later and both the deployment duration and the lead time come out wrong. - Change-failure event (
POST /api/v2/dora/incident) when the deploy is a hotfix. This feeds change failure rate and time to restore. Because production deploys are triggered by a push tomain, the ref is alwaysmain, so a hotfix is recognised from the merge commit GitHub writes for ahotfix/*PR (Merge pull request #N from <org>/hotfix/...). That is reliable because PRs intomainmay only come fromrelease/*orhotfix/*and always land as a merge commit, never a squash. See Branching & releases.
The failure window is reported as previous production commit to hotfix deployed: the incident
started_at is the commit time of HEAD^ (the version that was live and broken) and
finished_at is the moment the hotfix finished deploying. That is an approximation, because it
assumes the breakage arrived with the previous production deploy, and it is the best signal CI
has today. If HEAD^ cannot be resolved (a shallower checkout than fetch-depth: 2) the start
falls back to the deploy start rather than to finished_at, so a failed lookup cannot report a
zero-length incident and flatter a time-to-restore figure.
Robustness¶
Emission is best-effort and can never break a release:
.github/actions/dora-event/dora-event.shnever exits non-zero. A missing key, a Datadog outage, a 4xx or a malformed input logs a GitHub warning annotation and returns 0.- The calling step is additionally guarded with
continue-on-error: true. - The steps are gated on the pipeline outcome the
finishjob resolves from every other job, so nothing is emitted for a run that failed or was cancelled anywhere along the way. - Missing
DATADOG_API_KEYskips emission with a warning, matching how the sourcemap upload steps behave.
No new repository secret is needed: the action authenticates with the existing
DATADOG_API_KEY, the same secret the sourcemap and Cloud Run instrumentation steps already
use. The DORA intake authenticates with the API key alone, so no application key is sent.
Still to do by hand¶
The event pipeline is only half the deliverable. These parts live in the Datadog console and on Confluence and are not in this repo:
- Keep every service's DORA source set to API. Datadog can also derive deployments from
APM version tags, and a service with both sources on counts each release twice. That is what
PPL-3419 found:
clinicians-app-frontendhad APM deployment tracking on, so it produced a duplicatesource:apm_deploymentsevent for every release, plus a third for the web bundle, whose RUM version carries theBUILD_IDsuffix and reads as a separate deployment. Check this whenever a DORA number looks inflated: filter the deployments list bysourceand every event should sayapi. - Confirm events are arriving. After the next production deploy, open Datadog → Software
Delivery → DORA Metrics and check exactly one deployment event exists for
perci-platform. Confirm the exact metric and event field names in the console before wiring widgets to them. - Build the dashboard. One dashboard, four widgets (deployment frequency, lead time for
change, change failure rate, time to restore) over
service:perci-platform. If Datadog offers an out-of-the-box DORA dashboard, clone that rather than hand-building it. - Record the baseline. Once a full four weeks of events have accumulated, write the starting value of each of the four metrics onto the engineering strategy page in Confluence, so later improvement is measured against something rather than asserted.
- Review on a cadence. Recommended: review the four numbers monthly at the engineering review, and re-baseline quarterly. Treat a metric that has not moved for two quarters as a signal that the underlying delivery problem was misidentified, not that the metric is wrong.
Known gaps¶
- Hotfixes are the only change-failure signal today. A production incident fixed by a normal release, or one resolved without a code change, does not produce an incident event, so change failure rate and time to restore are floors, not exact figures.
- Canary auto-rollback is not wired. That signal belongs to PPL-3149 and does not exist
yet. The seam is already in the action: call it with
emit_incident: 'true'plusincident_nameandincident_started_at(the rollback knows when the bad version went out, so it can report a far more accurate failure window than theHEAD^approximation above). - The pipeline's own
rollback-backendjob does not emit an incident event either. A verification failure that triggers an automatic Cloud Run rollback is invisible to change failure rate today. It uses the same seam as the item above. - Staging deploys emit nothing. All emission is gated on
inputs.environment == 'production', so there is no staging series to filter out. Anenv:stagingdeployment event in Datadog did not come from this pipeline, and is a sign that a service has an APM-derived DORA source switched on. - There is no per-service breakdown any more. Deployment frequency, lead time, change failure rate and time to restore all describe the release train, not the backend or either app individually. Nothing is lost today, because the components never ship apart. If they ever do, this is the first thing to revisit.
Troubleshooting¶
release must be cut from the exact origin/develop commit¶
version must be greater¶
Pick the next semver above the version currently on both app pubspecs. Do not reuse a version that has already shipped.
remote branch already exists: origin/release/X.Y.Z¶
An earlier cut of the same version left its branch behind, usually because the release PR was closed instead of merged. Delete the stale branch and cut again:
If the release PR is still open, do not recut. Merge develop into the release branch instead
(see QA fixes).
PR branch appears to be grounded in main¶
Close the bad release PR and re-run pnpm run cut-release -- X.Y.Z from the latest
origin/develop.
.gitignore files were added¶
Remove generated build output from the release branch. Common examples are Flutter Linux
generated files under apps/*/linux/**, .dart_tool, and build/.
mcp/*/package-lock.json¶
Remove the npm lockfile. MCP packages use pnpm, and the root pnpm-lock.yaml is the
lockfile source of truth.
Back-merge PR has conflicts¶
Check out sync/main-to-develop, merge origin/main, resolve conflicts, then push the
branch. Keep the PR as a merge commit into develop.