Testing & the CI/CD release gate¶
This page is the source of truth for how the Flutter apps (perci-platform-members
and perci-platform-clinicians) are tested and what must be green before a change
can be released. The goal is continuous delivery: a change merges and releases as
soon as the gate is green, multiple times a day. The backend stack has its own gate and
its own coverage flags - see Backend coverage at the end of this page.
Where QA testing happens (feature-branch first)¶
We work in feature branches: a long-lived epic branch is cut at the epic level, and each piece of implementation branches off it and PRs back into the epic branch (see Epic (feature) branches). Opening a PR with the preview label spawns a preview environment (members, clinicians, backend) for that branch.
Test the feature in its own branch's preview, before it merges. This is the in Test
stage of the Delivery workflow; do not wait until the change is on
develop / staging to verify it. By the time work reaches On Staging, that step is for
regression against everything else, not first-pass feature testing.
The only exception: infrastructure that can't run in a preview
If a change depends on infrastructure that genuinely cannot be stood up in the feature branch's preview (for example a new queue, cron, external webhook, or environment-specific config), test it on staging after merge. This is the exception, not the norm, so call it out on the ticket to explain why it skipped feature-branch testing.
Test types¶
Flutter apps¶
Both apps use the same four layers, and the same tooling (patrol, golden_toolkit,
mockito).
| Layer | Tool | Lives in | Runs in CI |
|---|---|---|---|
| Unit (domain/data) | flutter_test |
test/features/**/domain, **/data |
every PR (blocking) |
| Widget | flutter_test + ProviderScope overrides |
test/features/**/presentation |
every PR (blocking) |
| Golden | golden_toolkit (@Tags(['golden'])) |
next to the widget, in goldens/ |
every PR (blocking) |
| E2E | patrol (Chrome/web) |
patrol_test/ |
pre-release / on-label (separate) |
Backend (functions)¶
| Layer | Tool | Lives in | Runs in CI |
|---|---|---|---|
| Unit / integration | vitest |
functions/src/**/*.test.ts |
every PR (blocking) |
| Firestore security rules | vitest + @firebase/rules-unit-testing + Firestore emulator |
functions/src/__tests__/firestore-rules/ |
every PR (blocking) |
Run the rules suite locally:
The suite covers member self-access, clinician self-access (notifications & calendar),
cross-UID isolation (another member's or clinician's documents are denied), the
users-public point-read-only policy, fully locked collections (plans, logs), and
deny-by-default for unauthenticated clients. The rules and the config that drives the
emulator both live in apps/perci-platform-backend/ (firestore.rules, firebase.json).
Mocking convention¶
- Prefer hand-rolled fakes that implement the domain repository interface for unit/provider/widget tests - they are explicit, fast, and need no codegen.
- Use
mockito(@GenerateMocks+build_runner) for new tests where call verification or mocking a concrete SDK/generated client adds real value. - Providers are tested with a
ProviderContainer(orProviderScope) overriding the repository/datasource provider with a fake. ForautoDisposeproviders, hold a listener before awaiting so the provider is not disposed mid-load.
The regression ratchet¶
A production regression closes with the test that would have caught it, at the lowest layer that can express the failure:
| The bug is in… | Test it with |
|---|---|
| A pure function, mapper, validator, or domain rule | Unit |
| A request/response shape against a generated client or backend contract | Unit against the client (a contract test), not a live call |
| Provider/notifier state, or a widget's reaction to state | Widget (+ ProviderScope overrides) |
| Layout, sizing, or structure that silently collapsed | Golden |
| Only the whole-flow interaction: navigation, auth, real cross-screen wiring | E2E (patrol) |
Pick the cheapest layer that actually fails on the bug. Reaching for E2E because it is quicker to write buys a slow test that catches less and flakes more; reaching for a unit test when the defect was the wiring between two screens proves nothing. The layer is a claim about where the defect lives, so state it on the ticket.
Backend regressions follow the same rule with vitest: a unit test over the handler or
service, or a route-level test against the Zod schema, before anything that needs emulators.
The process around this is in Delivery. This section answers only the "which layer?" question.
Shared harness¶
packages/perci_platform_test_shared is the single source of truth for the
fiddly, app-agnostic test setup: the silent network-image HTTP layer, the Firebase
Analytics fake, the package_info / secure_storage / datadog channel mocks,
golden_toolkit configuration, device presets and the Firebase core mocks. Each
app keeps a thin test/.../golden_harness.dart that calls
GoldenHarnessBase.baseGlobalSetUp() then wires its own Firebase init, auth
manager and FFAppState. Patrol widget wrappers stay per-app (they embed each
app's root widget).
The release gate (.github/workflows/ci-flutter.yml)¶
The two required PR status checks are Flutter checks passed (from
ci-flutter.yml) and Backend checks passed (from ci-backend.yml).
Both names are pinned in the branch rulesets — the workflow files can be
renamed, the aggregate job names must not be.
A PR to develop or main must pass:
- Code generation -
melos run build_runner(openapi + freezed + riverpod). - Analyze (errors block) -
flutter analyze --no-fatal-infos --no-fatal-warnings. Errors fail the build. Warnings/infos are reported but not yet fatal - they are a ratcheting backlog (see below). Flip to fatal-warnings once the count hits zero. - Unit + widget tests -
melos run test(excludes goldens), with coverage. - Coverage threshold - total line coverage must be
>= FLUTTER_MIN_COVERAGE(a repo variable). Ratchet this up toward 80%; never lower it. - Goldens - a separate blocking job (see Goldens).
Patrol E2E runs in a separate workflow, not on every PR (see E2E).
Alongside the blocking checks, a non-blocking Code complexity job runs
dart_code_linter over the app(s)
in scope for the PR (members, clinicians, or both - the same scoping the other
jobs use).
The metric thresholds (cyclomatic complexity, maximum nesting level, number of
parameters, source lines of code) live in each app's analysis_options.yaml.
The tool is deliberately not registered as an analyzer plugin: a legacy plugin
re-analyses the whole app inside the IDE's analysis server and splits the pub
workspace into extra analysis contexts, which is what made in-IDE highlighting
take 10-15 s (see "Analyzer performance" in running-locally.md).
Metric violations in the files a PR touches appear as inline PR annotations
(via scripts/frontend/dcl-github-annotations.dart — the tool's own GitHub reporter
does not cover metrics); the full per-app reports are uploaded as the
complexity-reports artifact, and when a PR's own files cross a threshold the job also maintains a sticky PR comment listing those findings - the detail a future blocking check would fail on. The job never blocks a merge. Run the same analysis locally with melos run complexity.
Coverage baseline & ratchet¶
Set the repo variable FLUTTER_MIN_COVERAGE to the current measured floor, then
raise it as the backlog is burned down. The immediate purpose of the gate is
non-regression (coverage may not drop); the long-term target is 80% on
hand-written code, reached by ratcheting.
The floor follows the coverage actually present on the target branch. Raise it after each coverage-improving PR; never lower it below the branch's real coverage.
The merged report excludes generated Dart - *.g.dart, *.gr.dart,
*.freezed.dart, *.mocks.dart and **/generated/** - so the figure the gate
enforces and the figure Datadog reports measure the same thing. Both come from
scripts/merge-flutter-coverage.sh, shared by ci-flutter.yml and
coverage-develop.yml. On the PPL-3430 branch the exclusion removes 671 files
and 13,155 lines.
| Scope | Line coverage |
|---|---|
| Merged, generated excluded (what the gate checks) | 46.73% (38119/81569) |
| Merged, raw (for comparison with older figures) | 42.04% (39870/94845) |
| perci-platform-members | 46.75% (19543/41803) |
| perci-platform-clinicians | 37.96% (14359/37829) |
| perci-platform-frontend-shared | 39.23% (5968/15213) |
PPL-3430 moved the merged figure from 25.97% to 46.73%. Set
FLUTTER_MIN_COVERAGE to 46 once it is on develop.
perci-platform-frontend-shared alone has since moved from 39.23% to
80.94% (12600/15568), measured the same way (flutter test --coverage in
the package, generated Dart excluded). The denominator grew from 15,213 to
15,568 because the new tests load libraries nothing loaded before - the effect
the warning below describes. Re-measure the merged figure before moving
FLUTTER_MIN_COVERAGE; the other two packages are unchanged by this work.
The denominator moves
flutter test --coverage only instruments libraries the tests actually
load, so adding a test for a previously-untouched file adds its lines to
the denominator as well as the numerator. A shallow smoke test over a large
widget can therefore lower the reported total. Cover a chunk properly
rather than merely touching it, and re-measure the whole suite - never a
subset - before quoting a number.
The reported percentage is not coverage of the codebase¶
That mechanic has a bigger consequence than a wobbly number. A file no test loads appears in the report not at all - it is missing from the numerator and the denominator, so it is invisible rather than counted as zero. On the PPL-3430 branch 972 of 2,105 non-generated Dart files (46%) were absent entirely. The percentage is therefore coverage of the subset the suite happens to import, and it flatters: the files nobody has tested are exactly the ones excluded from the sum.
The fix is a library that imports everything, so the denominator is the whole
package regardless of what the tests exercise. perci-platform-frontend-shared
has one at
test/coverage/all_libraries.dart;
adding it moved that package from 37.27% over 200 files to 28.21% over 292
with the numerator unchanged. Lower, and true - and the shared figure in the
table above is measured on that honest denominator, while the two app figures
still are not.
Neither app can have one yet.
lib/custom_code/actions/get_current_used_devices.dart in the shared package
calls MediaStreamTrack.getSettings, which dart_webrtc provides on web but
not on the Dart VM the test runner uses, and it is re-exported from the
custom_code/actions barrel - so it is dragged into any complete compile of
either app. Until it is made VM-safe or the barrel is split, the two app figures
remain subset figures, and the merged number with them.
Datadog coverage reporting¶
Every suite uploads its merged lcov report to Datadog Code Coverage (EU site):
- PRs - the
datadog-ci coverage uploadstep inci-flutter.ymlandci-backend.yml. On apull_requestevent Datadog attributes the upload toGITHUB_HEAD_REF, so these only ever produce feature-branch coverage. - develop -
.github/workflows/coverage-develop.ymlre-runs the suites on push todevelopand uploads, giving Datadog the default-branch baseline it diffs PR coverage against. Without it thedevelopbranch view is empty.
The develop workflow deliberately runs tests only (no lint/build/threshold gate, no PR comment) and does not reuse the required check names, so it can never block a merge. Its env scaffolding mirrors the PR jobs; keep the two in sync or the baseline stops being comparable.
Every upload carries a flag - flutter, backend-functions or
backend-web-professionals - set on the
DataDog/coverage-upload-github-action step. Flags are what
code-coverage.datadog.yml
scopes carryforward to, and what the coverage gates will be scoped to. Most PRs
touch one side of the repo only, so without them a backend-only PR reads as if all
Flutter coverage had vanished. An unflagged upload merges into one
undifferentiated pile that cannot be gated per stack. If a new package starts
uploading, give it its own flag rather than reusing one of these.
code-coverage.datadog.yml (repo root) also maps paths to the Datadog services
they already report under, so coverage lines up with APM, RUM and DORA, and ignores
generated Dart output.
Datadog PR Gates¶
code-coverage.datadog.yml defines two
PR Gates, both
on total coverage. Patch coverage is measured in CI instead - see
Patch coverage in CI.
| Gate | Scope | Threshold | Effect |
|---|---|---|---|
total_coverage_percentage |
flag flutter |
20 | Fails the PR if Flutter total coverage drops below 20% |
total_coverage_percentage |
flag backend-functions |
80 | Fails the PR if backend total coverage drops below 80% |
threshold only accepts a number (0-100); there is no relative or auto value
in the config file. A non-numeric threshold does not just disable that gate - it
makes the whole file invalid, so services, ignore and carryforward stop
applying too and the UI reports the config as invalid. The two totals above are
floors set just under the current develop figures (backend 82.2%, Flutter 23.2%);
raise them as coverage climbs, or express non-regression as a rule in the
Datadog UI, which the file cannot do.
The totals are scoped by flag deliberately. Most PRs touch one side of the repo,
and carryforward
supplies the other side's figure from an ancestor commit. Scoping to the
services instead would evaluate each one separately and fail PRs over packages
the change never touched.
Scope multiplies evaluations, it does not filter them¶
services, codeowners and flags are all optional on a gate, and
every scope listed is evaluated independently -
"the gate does not combine coverage across services or code owners". Scope
therefore decides how many times a gate runs, not whether it applies to the
change in front of it. A scope covering code the PR never touched has no patch to
measure, and:
Coverage is -1 when Datadog has no data - a no data sentinel, not 0% - and a
gate reading -1 fails.
That is why there is no Datadog patch gate. Every scoping was tried and every one fails a PR that changes no coverable line, because there is genuinely nothing to measure:
| Patch gate scope | Result on a PR with no coverable changes |
|---|---|
flags: ['*'] |
#1752 (a pubspec.yaml bump) - one red check |
services: ['*'] |
#1896 (this file plus a doc) - four red checks, one per service |
| none | #1898 - one red check |
Scope changes how many times the gate goes red, never whether it applies. A gate that cannot be made blocking without making config-only PRs unmergeable is not a gate, so patch coverage moved to CI, where an empty diff can be skipped explicitly - see Patch coverage in CI. Do not re-add it here; the totals are the only thing this file gates.
Note that a gate is evaluated even when a PR's own CI skipped the test jobs and uploaded nothing - #1896's Flutter and backend aggregates passed off skips, and it was still gated. Path-filtering CI does not opt a PR out.
Datadog reports one status check per gate type, not per gate, so the two total
gates share a single Total coverage percentage check. A check is only blocking
once its context is added to the PR checks ruleset - a gate can fail while the
PR stays mergeable until then, which is the case today for both. Gates also never
apply retroactively: an open PR is only evaluated once a new commit is pushed
to it. A stale Patch coverage percentage check may linger on PRs opened before
that gate was removed; it is inert and disappears on the next push.
Gates can additionally be defined in the Datadog UI. Where the UI and this file both cover a scope, a PR must satisfy every threshold, so prefer keeping them here where they are reviewable alongside the code.
These sit alongside - not instead of - the in-CI gates: the vitest
coverage.thresholds and the FLUTTER_MIN_COVERAGE step above still run, and
still block. Both sets are floors, so keep them in step: the Datadog thresholds
are evaluated on the merged report, the in-CI ones per package before upload.
Patch coverage in CI¶
scripts/patch-coverage.mjs replaces the Datadog patch gate, which could not be
made blocking. It runs as a Patch coverage step in the Flutter and backend test
jobs, after the lcov report is written and before it is uploaded, and:
- reads the lines the PR adds, from
git diff --unified=0 HEAD^1 HEAD(HEAD^1is the base branch tip on apull_requestmerge-ref checkout, so the existingfetch-depth: 2is enough history); - keeps only those that appear as executable (
DA:) lines in the lcov report; - skips with a pass when that leaves nothing - the config-only case;
- otherwise fails below
PATCH_MIN_COVERAGE(default 80).
Added lines in files with no lcov record - a YAML file, a Markdown doc, a package the suite does not instrument - are ignored rather than counted as missed. That is the whole difference from the Datadog gate, and the reason this one can be required.
It is warn-only until the PATCH_COVERAGE_ENFORCE repository variable is set
to true; below-threshold PRs log a ::warning:: and pass. Flip the variable to
make it blocking - no code change, and it is already inside the required
Flutter checks passed / Backend checks passed aggregates, so nothing needs
adding to the ruleset.
The script is plain dependency-free ESM rather than TypeScript because the Flutter
job sets up no Node toolchain and runs it on the runner's stock node. Its unit
tests are TypeScript, in scripts/__tests__/patch-coverage.test.ts, and run with
the rest of the repo tooling via pnpm run test:release-tooling. Run it locally
against a report you have already generated:
Coverage paths must be repo-relative. Datadog files a report by the paths in
its SF: lines, so a package-relative report lands in a phantom top-level folder
in the File Explorer. Flutter is handled by the merge step, which prefixes each
package dir. Backend needs the same treatment: vitest runs inside the package and
emits SF:src/..., so every backend job rewrites the paths into
coverage/datadog/lcov.info before uploading (datadog-ci has no path-prefix
flag), leaving the original lcov.info untouched. It is not only cosmetic -
src/helpers/sleepSeconds.ts exists in both functions and web-professionals,
so unprefixed the two reports collide on one phantom path. Any new package that
starts uploading coverage needs the same rewrite.
Flaky test reporting¶
Datadog classifies a test as flaky when it sees both a passing and a failing
status for the same commit. Datadog has no native Dart/Flutter tracer, so its
Auto Test Retries and Early Flake Detection cannot run here - the JUnit XML
upload in ci-flutter.yml is the only ingestion path, and the retry has to
happen in the test run itself.
Each Flutter package that has tests therefore sets retry: 1 in its
dart_test.yaml. A failing test gets a second attempt, so a one-off flake no
longer reds the build and, more importantly, the pass/fail pair Datadog needs
actually exists.
That alone is not enough. flutter test's JSON reporter collapses the
attempts into a single entry - one testDone with result: success plus the
error event from the attempt that failed - and junitreport:tojunit turns
that into one <testcase> carrying a <failure>. Left as-is, a test that
recovered on retry would be uploaded as a hard failure and never as a pass, so
Datadog still could not classify it (and the data would be wrong).
scripts/frontend/flaky-junit-report.dart
runs between the conversion and the upload and puts the missing attempt back. It
reads the JSON report, which is the only source that distinguishes the two cases:
| JSON result | error events |
Meaning | XML after the script |
|---|---|---|---|
success |
>= 1 | recovered on retry | failing and passing testcase |
failure |
one per attempt | failed every attempt | failing testcase, untouched |
success |
none | passed first time | untouched |
Datadog then sees a pass and a fail for the same test on the same commit and marks it flaky, which feeds the repo's existing auto-quarantine and auto-disable policies. Genuinely failing tests are left alone and still fail the build.
The E2E lane solves the same problem separately: Patrol runs with
web-retries: 2 and reports a flaky-count (see The web lanes).
Analyze warning ratchet¶
melos analyze currently reports ~360 warnings/infos across the workspace, almost
all pre-existing in legacy FlutterFlow code (perci_library_9rk85z) and a few in
older test infra. There are no error-severity issues, so the error-only gate
passes today. Burn the warning count down (a chunk is auto-fixable via
dart fix --apply), then make warnings fatal in the gate.
Goldens¶
Golden tests run on every PR for both apps via flutter test --tags golden (the
golden job in ci-flutter.yml). The job is blocking: a failing golden
fails the required Flutter checks passed gate and prevents merge.
Cross-platform: render in boxes, not real fonts¶
Real fonts rasterise differently per OS (Windows DirectWrite, macOS CoreText, Linux
FreeType), so golden PNGs made on one machine never match another - a mixed Win/Mac/Linux
team plus Linux CI can't share real-font baselines. So our goldens do not load real
fonts: the harness never calls loadAppFonts(), so Flutter's test environment renders
all text in the Ahem font (every glyph a fixed square). Ahem output is identical on
every platform, so text no longer moves a baseline between machines. (This is the same
trick as Alchemist's "CI mode"; we do it directly rather than add the dependency.)
Ahem fixes text, not everything. Skia still rasterises shapes, borders and scaled
images slightly differently across CPU architectures, so a baseline is not bit-identical
between a dev machine and CI - see the rationale in each test/flutter_test_config.dart,
which measured ~1.7% on Apple Silicon vs Linux x64 and set the comparator tolerance to 3%
to clear it (8% in perci-platform-frontend-shared, whose goldens are more fill-heavy).
Consequence: goldens verify layout, sizing, colour and structure - not readable text (text shows as boxes). Text content is asserted by widget tests.
Linux CI is the authority, not your machine. A fill-heavy golden can sit just over the
local tolerance and still be green on CI: as of 13 Aug 2026 twelve members goldens fail on
macOS (plan_dashboard, mfa_security_page, mfa_setup_page, welcome_configuration,
read_only_*, home_dashboard) while being green on the Linux runner, where the golden
job gates every merge into develop. Before concluding your change broke a golden, run the
suite with and without it - if the same set fails both ways, it is host drift. Regenerate
baselines through the Flutter - Update Goldens workflow (which runs on CI) rather than
locally, or you will bake macOS rasterisation into the baseline.
- Regenerate through the Flutter - Update Goldens workflow, which runs on the Linux CI
runner and opens a PR.
flutter test <path> --tags golden --update-goldenslocally is fine for iterating, but do not commit the result from a non-Linux machine: text is platform-independent, the rest is not, so you would trade your own drift for everyone else's. - Not golden-tested (non-deterministic regardless of fonts): live camera/video widgets
(the old
meeting_room/waiting_roomgoldens) and the animatedWelcomePagewere dropped; a couple of ultra-narrow scenarios areskip-ed where the wider Ahem glyphs tip a flex-less row into overflow (covered by their wider siblings + widget tests).
E2E (patrol)¶
Patrol E2E runs on four lanes, all selecting tests by tag:
| Lane | Workflow | When | Backend | Gate? |
|---|---|---|---|---|
| Release candidate (web) | e2e-web.yml |
every PR into main or a release/* branch |
the candidate's own preview backend | fails its own run; not a required check |
| E2E authoring (web) | e2e-web.yml |
any PR that edits patrol_test/, the shared harness, or the lane's CI definition |
real staging (the apps' committed endpoints) | fails its own run; not a required check |
| Nightly (web) | e2e-web-nightly.yml |
02:30 UTC nightly + manual | real staging | advisory, deliberately not a required check |
| Native (Android) | e2e-native-manual.yml |
every PR into main (automatic) + manual dispatch, any app, any branch |
real staging via Firebase Test Lab | Native checks passed is a required check on main |
Tag vocabulary¶
Lane selection is tag-driven. The vocabulary lives in one place —
TestTags in packages/integration_test_shared — and raw-string vocabulary
tags are rejected by a lint step in ci-flutter.yml (a typo'd raw tag would
silently drop a test out of its lane). Zephyr case ids ('PPL-TNN') stay raw.
| Tag | Meaning | Selected by |
|---|---|---|
webSafe |
Honest-green, web-runnable, standalone (no seed step beyond what the lane's own action provides, no native device, non-destructive on shared staging) | e2e-web.yml and e2e-web-nightly.yml (--tags webSafe) |
native |
Needs a real Android device (Stripe CardField, WebRTC, permissions, file_picker) | e2e-native-manual.yml — automatically on PRs into main, or via target / tags on manual dispatch |
smoke / regression / desktop / tablet / mobile |
Suite descriptors; selectable ad hoc via the manual lanes' tags inputs |
— |
Add webSafe to a case only once it is confirmed green and standalone;
seed-dependent cases join when a seed pre-step lands. patrol bakes --tags
into the generated bundle at build time, so each lane's build compiles exactly
the matching cases.
The web lanes¶
Both apps share e2e-web.yml. It runs against Chrome (web) via Playwright and
takes ~75 minutes, so it does not run on every PR — it has two lanes:
- Release candidate — every pull request into
mainor arelease/*branch, both apps unconditionally. A release ships both, so there is nothing to gain from filtering the matrix by which app's files changed. - E2E authoring — any pull request that edits the Patrol tests
(
apps/*/patrol_test/), their shared harness (packages/integration_test_shared/) or the lane's own CI definition, so a test change is proven by running it. Only the app(s) whose tests changed run; a shared or CI change runs both.
Every other pull request has no E2E lane: they gate on the unit, widget and contract suites, and the nightly run below is the net between releases.
A release candidate is tested against its own preview backend, not shared
staging. The preview hosts are deterministic from the PR number, so this lane
computes backend-pr-<N>.perci.dev itself and rewrites the app's
environment.json — no handover from preview-pr.yml, which used to call this
workflow (workflow_call) and tie every preview deploy to a 75-minute suite.
What it does still need from preview-pr.yml is when that backend is ready, so
the Wait for the preview backend step polls that workflow's backend job for the
PR's head sha. Only the job's conclusion separates the three outcomes:
success is a fresh backend; skipped means the content-hash gate found the
backend unchanged since the last push, so the preview host is already serving a
content-identical build; failure means the host is still answering from the
previous revision. The last case fails the E2E run rather than quietly falling
back to staging, which would hand a release candidate a false green. That is also
why the signal cannot be a commit-exact /_health probe: on a skipped deploy
the preview backend legitimately reports an older sha. The job name is pinned in
this workflow's PREVIEW_BACKEND_JOB_NAME env — renaming that job in
preview-pr.yml without updating it turns the wait into a hard failure, deliberately.
The E2E-authoring lane has no preview backend to wait for and tests the staging
endpoints committed in each app's environment.json.
The nightly lane (e2e-web-nightly.yml) runs the same webSafe subset for
both apps against real staging, on the 02:30 UTC cron or on demand. It is
deliberately not wired to pull requests: a PR into main or release/*
already runs that subset through the lane above, and that lane is the one that
publishes the Zephyr cycle. Triggering both ran every case twice per release
push. The nightly lane keeps one thing the release-candidate lane cannot have —
it uploads its JUnit report to Datadog Test Optimization, which needs a Datadog
key that e2e-web.yml is deliberately never handed, because that lane builds
and runs PR-authored code.
SSO sign-up: the identity-provider sidecar¶
The member SSO sign-up case (sso_sign_up_test.dart, Zephyr T4) is the one web
case that cannot stay inside the Flutter page: the real flow leaves the app for
the identity provider, and the Descope exchange code the provider hands back is
single use and expires within minutes, long before a web bundle finishes
building. So the identity leg runs in a sidecar the test calls at run time.
patrol_test/tools/sso/serve_sso_mint.mjs enrols a throwaway account on the
team's test identity provider (Authentik at auth.perci.dev, whose enrolment
flow is public, so no credentials or secrets are involved), performs the
real login and consent headlessly, captures the code from the redirect without
loading it, and returns it with the identity the app should prefill. The test
then enters at /sso-callback?code=…, the exact route the browser redirect
uses, and drives the sign-up forms to the dashboard, asserting the identity the
provider supplied is what the app prefilled.
.github/actions/patrol-web-test starts the sidecar for the members app and
passes --dart-define=SSO_MINT_URL=…; without that define the case is skipped,
never failed. Locally, zsh patrol_test/tools/run_sso_patrol.sh does the same.
The sidecar mints against the backend the app's environment.json points at,
so a release candidate exchanges codes its own preview backend produced. Each
run leaves one sso-e2e-<epoch> Authentik user and one staging member behind,
the same footprint as the voucher sign-up case. Details, overrides and hygiene
notes: apps/perci-platform-members/patrol_test/tools/sso/README.md.
Zephyr: the automated release cycle¶
QA plans and records release testing in Zephyr Scale, so the automated coverage has to land
there too rather than living only in a CI log. On a release/* or hotfix/* PR into main,
the zephyr job in e2e-web.yml publishes the web lane's results as Zephyr test cycles, one per
app, role and device, with each case marked Pass or Fail:
That mirrors the manual cycles (Release 1.1.23 - Clinician App - Nurse Admin - Manual): app
first, then role where one applies, then device, with the execution mode last.
The app comes from the artefact the report arrived in (members-test-logs /
clinicians-test-logs), since a spec title does not carry it. The device comes from the lane:
the Chrome lane is a single 1920x1080 viewport so it publishes Desktop, and Firebase Test Lab
runs real devices so it will publish Mobile (PPL-3444). The existing desktop/tablet tags are
not used for this, because they describe intent rather than the viewport that actually ran.
Every cycle is attached to the release's Zephyr Test Plan, named Release <x.y.z> to match
the plans QA already keeps (Release 1.1.20 upward). The plan is created if the release does not
have one yet, so the automated cycles hang off the same plan as the manual ones.
Worth knowing if you touch that code: the link is written from the plan side
(POST /testplans/{planKey}/links/testcycles) and the field is testCycleIdOrKey, even though
the payloads report the relationship as testCycleId and a cycle reads its own links back from
GET /testcycles/{key}/links. The published API docs are unreachable, so both facts came from the
live API; the obvious reading of the payloads gives the wrong side and the wrong field name.
If the plan cannot be created or linked, the cycles and executions publish anyway and the run warns. A plan groups the record, whereas the executions are the record.
The role comes from the test's own TestTags.role* tag, and this is the part worth
understanding. A Zephyr case is labelled for every role a manual tester covers with it: PPL-T47
carries Nurse, Admin_Nurse and AHP. An automated test signs in as one. Grouping on the
case's labels would therefore record a Pass in the AHP cycle for a run that never signed in as
AHP, so instead each test declares the role it actually runs as and only ever lands in that role's
cycle. A test that genuinely signs in as two roles carries both tags and appears in both.
Tag the role the test really exercises:
patrolTest(
'[PPL-T37] Appointments: clinician views their appointment list',
tags: [TestTags.regression, TestTags.webSafe, 'PPL-T37', TestTags.roleNurse],
For a test parameterised over CliniciansRoles.roles, use TestTags.roleTagFor(role) so each
generated test carries only its own role. A case that signs in as nobody (the member app, or
PPL-T10 which only asserts the pre-auth sign-in UI) carries no role tag and lands in the app's
roleless cycle.
Note that the role names are not covered by the raw-string tag lint in ci-flutter.yml. Unlike
the lane vocabulary, 'Nurse' is also a legitimate credentialRegistry['Nurse'] key, so a
single-line grep cannot tell a stray raw tag from a login call.
The link between a Patrol test and its Zephyr case is the case id in the test name:
patrolTest(
'[PPL-T25] Member details Message button opens the conversation',
tags: [TestTags.regression, TestTags.desktop, 'PPL-T25'],
patrol's web runner uses the Dart test name as the Playwright spec title, so
scripts/zephyr/publish-e2e-results.ts reads the case ids straight out of the
playwright-report/results.json the Patrol job already uploads. Both authored shapes are
understood — [PPL-T196] [PPL-T187] and the combined [PPL-T47 T48 T49]. A case with no id
in its title is invisible to Zephyr, which is the main thing to get right when adding a test.
Each execution also carries:
- The measured run time, in Zephyr's "Actual Time" field, taken from the attempt that decided the outcome. For a case that passed on retry that is the successful attempt, so the figure is the test's own run time rather than the sum of its failed attempts.
- The failure message in the execution comment, with Playwright's ANSI colouring stripped and the text escaped and truncated so a long assertion diff cannot break the rendering.
- An evidence pointer, on every execution. Pass or fail, the comment names the CI artefact
(
clinicians-test-logs), links the run, and says to openplaywright-report/index.html. A failure additionally quotes its video path. Reviewing a failure usually means comparing it against a passing run, so an execution with no pointer at all just forces a hunt through the Actions UI.
Playwright records attachment paths as absolute paths on the machine that ran the suite, so the
video path is re-anchored onto the downloaded artefact by its test-results/ segment, and is
only quoted once the file has actually been found.
Evidence is referenced, not attached. Zephyr Scale's v2 API rejects
POST /testexecutions/{id}/attachments with 405 Method Not Allowed, so nothing can be uploaded
from CI. The reference is more useful anyway: the artefact holds the HTML report, every video and
the full log, not just one file.
Because those references are the release's evidence and QA reviews the cycle after the release
ships, a release or hotfix candidate keeps its artefacts for 90 days instead of the ordinary
7 (see retention-days in e2e-web.yml). At 7 days every citation in Zephyr would rot within
the week.
Playwright traces are not produced today: patrol's playwright.config.ts sets video but
never trace, and neither its env vars nor patrol_cli expose a way to turn it on. Enabling
them needs a change upstream in patrol.
Behaviour worth knowing:
- A case that passed only on retry is recorded as
Passwith a flaky note, matching the lane's own verdict; a case that failed every attempt is aFail. - A combined case (
[PPL-T47 T48 T49]) attaches the same video to each of its executions, since each case is a separate row and each should carry its own evidence. - Skipped cases get no execution at all — the run says nothing about them, and a row implying otherwise is worse than no row.
- Re-running the workflow reuses the existing cycle for that version instead of creating a duplicate.
- Publishing is not a gate. It never fails the E2E verdict, and a missing
ZEPHYR_API_TOKENrepository secret only logs a warning. - A case id present in the suite but unknown to Zephyr is reported in the job summary rather than dropped silently, so the two inventories can be reconciled.
The version comes from the PR head's apps/perci-platform-members/pubspec.yaml, not the branch
name, so hotfix branches (which carry no version in the branch name) name their cycle correctly.
The cycle is also linked to the Jira release version, so QA can filter cycles by release. The
lookup is by exact name and tries Platform <x.y.z> before a bare <x.y.z>, because PPL contains
both conventions plus unrelated Member App x.y.z and Clinical App x.y.z versions that would
make a loose semver match pick the wrong release. It needs the JIRA_BASE_URL, JIRA_USER_EMAIL
and JIRA_API_TOKEN secrets; without them, or when no version matches, the cycle is still
published, just unlinked.
Cases covered here should carry the automated label in Zephyr, which is what keeps
/zephyr-release-cycles from also pulling them into the manual cycles.
The native lane publishes too¶
e2e-native-manual.yml publishes into the same cycles, as Mobile. gcloud firebase test android
run only prints a leg-level pass or fail, but FTL writes a test_result_*.xml per device and shard
into its results bucket, with a <testcase> per Dart test. Each leg pins --results-dir so the
path is predictable, recovers the bucket name from FTL's own output, and gsutil cps the XML into
its artefact.
Patrol names each native test <filename> [PPL-T93] <title>, because its JUnit runner is
parameterised over the Dart test names, so the same key extraction works unchanged.
Two things differ from the web lane:
- JUnit XML carries no tags, so the app and role cannot come from the results. Each leg runs
pnpm run zephyr:case-metaagainst the commit under test and ships the resultingcase-meta.jsonbeside the XML; the publisher joins on it. That is also why the map is generated in the test job rather than the publish job, which checks out the base branch. - The device is derived from the lane, not a flag. FTL runs real devices, so its results are
always
Mobile. Deriving it per result means a single run covering both lanes cannot mislabel the native half, which one global--devicewould.
The evidence pointer differs too: an FTL artefact holds JUnit XML, logcat and the device video, so
a native comment points at the Firebase Test Lab results rather than a playwright-report that was
never there.
The XML shape is FTL's standard instrumentation output and is parsed tolerantly (both
<testsuites><testsuite> and a bare <testsuite>, failure or error, skipped), but it has not
yet been confirmed against a live run.
Dry-run the mapping locally against a downloaded report:
The native lane¶
Some flows cannot run in headless Chrome at all — Stripe's CardField is an
Android platform view, WebRTC and GetStream need real media permissions,
file_picker needs a real file system. Those cases are tagged native and run
on real Android devices in Firebase Test Lab via e2e-native-manual.yml.
Every PR into main — so every release and hotfix candidate — runs the native
matrix automatically, gated by the required Native checks passed check. The
matrix is over (app, device, selector), not just app, because the cases disagree
about the device they need:
| Leg | Device | Selects | Seeds |
|---|---|---|---|
| members purchase | MediumPhone.arm portrait |
--tags native (PPL-T100, Stripe CardField) |
seed_screening_purchase.sh, re-run per run because the pathway is consumed |
| clinicians video / in-call | MediumTablet.arm landscape |
--tags "native && !PPL-T140" |
seed_join_video_appointment.mjs |
| clinicians file picker | MediumTablet.arm landscape |
--target document_upload_native_ppl_t140_test.dart |
none; pushes the fixture PDF via --other-files |
The clinicians legs need a tablet in landscape because the member-details page is an unconditional two-column layout with no mobile breakpoint — on a phone viewport the right panel is crushed and its tap targets go off-screen.
Seeding now happens inside the workflow, so no manual pre-step is needed.
Manual dispatch still takes either app, any branch, a single target or a tags
selector, with optional FTL sharding for parallelism.
The call-member cluster¶
These four cases look alike and are routinely lumped together, but they belong on different lanes. Only one of them places a real call:
| Case | Places a real call? | Lane | Needs |
|---|---|---|---|
PPL-T91/T96 call_member_connect |
yes | web only | the auto-answer Twilio peer |
| PPL-T92 cancel modal | no, confirm is asserted but never tapped | web | a member with a phone |
| PPL-T93 mic denied | no | native only | the FTL harness withholding RECORD_AUDIO |
PPL-T97 call_member_tier1 |
no | web | a member with a phone |
PPL-T91/T96 cannot run on the native lane at all: the app never registers an
Android Telecom PhoneAccount, so the twilio_voice plugin rejects makeCall with
"No registered phone account". The clinician calling path is web-primary, and the
case is honest-green on the web lane because --use-fake-device-for-media-stream
lets the Twilio Voice JS SDK place the call and it genuinely connects to the
auto-answer peer. It is deliberately not webSafe — it places a real, billed PSTN
call, so it does not belong on a per-PR lane. See PPL-3379.
PPL-T93 is the mirror image: native-only, because the androidTest harness grants
RECORD_AUDIO per Dart test name and deliberately withholds it for this one case
(MainActivityTest.java), which is what makes the denied path reachable. It needs
no Twilio number and places no call, so it runs on the automatic native lane.
Every user story should land with an automated E2E that walks its flow, so the patrol suite
grows into a full regression net. That suite is the release regression gate: because it runs
on every PR into main or a release/* branch, a green run is what lets us ship confident there are
no regressions, without a manual re-test pass. Adding the E2E is part of the
Definition of Done.
The shared harness lives in patrol_test/ (patrol_setup.dart, patrol_widget_wrapper.dart,
helpers/clinician_session.dart, pages/). patrol test discovers patrol_test/
by default. Clinician flows covered: sign-in (+ forgot-password), sign-out, members-list
search, members-list filter, member details (contact + demographic), member medical
record (documents + screening sections), appointments (tabs), messages, payments,
and the Learn hub (open + search). Section/tab flows are permission-gated and skip cleanly when
a role lacks access. These are authored + analyze-verified; runtime execution is the
patrol CI job.
Autofix on failure¶
When a Patrol run fails, an agentic follow-up (.github/workflows/e2e-web-autofix.yml)
investigates instead of waiting for a human's slow edit-and-rerun loop. It decides
whether the failure is a flaky test (e.g. a waitUntilVisible hitting a refresh
animation) or a real functionality break — probing live UI state with marionette and
re-running only the failing target as the authoritative gate — then opens a draft PR with
the fix (test-side for a flake; a minimal product fix plus a report for a real break).
See Patrol autofix for the design.
Running locally¶
melos bootstrap # resolve deps (first time / after pubspec changes)
melos run build_runner # generate openapi + freezed + riverpod
melos run analyze # static analysis
melos run test # unit + widget (excludes goldens)
# single app, with coverage
cd apps/perci-platform-members && fvm flutter test --coverage --exclude-tags golden
# goldens for one file
fvm flutter test <path> --tags golden # compare
fvm flutter test <path> --tags golden --update-goldens # regenerate
Python: the SFTP referral importer¶
The only Python in the repo is the referral-import Cloud Function at
infrastructure/sftp-referrals/functions/process-referral/, which parses partner
referral PDFs and posts them to the EBP Referral API. It has its own harness
because it shares nothing with the Flutter or Node stacks.
| Layer | Tool | Lives in | Runs in CI |
|---|---|---|---|
| Unit | pytest |
test_*.py, beside the code |
every PR (blocking) |
cd infrastructure/sftp-referrals/functions/process-referral
pip install -r requirements-dev.txt
./run_tests.sh # what CI runs — do not use bare `pytest`, see below
Tests sit next to the module rather than in a tests/ subdirectory, so
from parsers.base import ... resolves without packaging the function — pytest
prepends the test file's directory to sys.path. requirements-dev.txt pulls in
the full runtime requirements on purpose: the parser imports pdfplumber at
module scope, and installing the real dependency set means CI catches an import
that only resolves on someone's laptop. That is the failure mode behind the exact
pin on rapidocr-onnxruntime (PPL-2200).
Individual files also run standalone (python3 test_slack_notify.py) via a
__main__ block. That is a convenience, not the contract — CI runs run_tests.sh.
One interpreter per test file¶
run_tests.sh invokes pytest once per test_*.py, and bare pytest over the
directory does not work. Every test file fakes the modules main.py imports
(requests, google.cloud.storage, functions_framework, parsers) by
installing stubs into sys.modules before import main.
In one process main is imported once, and it binds whichever stubs were
installed at that moment. The first file pytest collects wins; every later file's
stubs land in sys.modules too late to reach main. The symptom is a file whose
assertions see an empty capture list, because main is posting into another
file's fake — which reads like a bug in the code under test and is not one.
main.py also builds a module-level storage_client at import time, which no
amount of later rebinding can redirect.
Isolation is what these tests were written to expect, so the harness gives it to
them. The proper fix is a shared conftest.py installing one set of stubs for the
whole suite (PPL-3418); when that lands, this script can go.
If you add a test file that stubs a module main imports, keep the stub
module-level and do not rely on another file's fakes being visible.
Fixtures: never use real referral data¶
Referral PDFs are member clinical records. Names, policy numbers, claim numbers and email addresses in tests are always synthetic — do not paste values from production, from a bucket object name, or from a log line into a test file. Committed test data lives in git history forever.
Parser logic is therefore factored into pure string helpers that need no PDF at all, which is how the majority of cases are covered. If a test genuinely needs a PDF, generate a synthetic one at test time matching the form's table layout; do not commit a real referral, redacted or otherwise.
The gate¶
The job is SFTP referrals checks in ci-backend.yml, gated on a
sftp_referrals paths filter and wired into the Backend checks passed
aggregator. It reuses that existing required check rather than adding a new one,
so no branch-ruleset change is needed for it to block a merge.
Before this existed, a PR touching only infrastructure/sftp-referrals/** saw
Backend checks passed and Flutter checks passed both go green in nine seconds
off skipped jobs, with no test executed. If you add another stack outside
apps/**, give it a filter and add it to an aggregator's needs — a required
check that passes because everything under it was skipped is worse than no check,
because it reads as assurance.
Coverage backlog (path to "all functionality covered")¶
Comprehensive coverage is delivered by ratcheting FLUTTER_MIN_COVERAGE, not in a
single pass. Work is chunked by page or feature, one chunk per PR, each measured
against a full-suite run.
Regenerate this ranking from a merged report at any time:
awk -F: '/^SF:/{f=substr($0,4);h=0;n=0} /^DA:/{split(substr($0,4),a,",");n++;if(a[2]+0>0)h++} /^end_of_record/{if(n>0)printf "%6d %6d %5.1f%% %s\n", n-h, n, 100*h/n, f}' coverage/lcov.info | sort -rn | head -40
Largest remaining debts¶
Measured 13 Aug 2026, after the PPL-3430 chunks. Uncovered lines first, since those are lines already in the denominator - covering them raises the total without moving the goalposts.
| Uncovered | Total | Cov | File |
|---|---|---|---|
| 956 | 956 | 0.0% | clinicians new_a_p_p/calendar/pages/calendar_home/calendar_home_widget.dart |
| 591 | 856 | 31.0% | shared components/address_field_library/address_field_library_widget.dart |
| 511 | 948 | 46.1% | clinicians new_a_p_p/your_details/personal_details/personal_details_widget.dart |
| 486 | 905 | 46.3% | clinicians new_a_p_p/appointment/appointments_page/appointments_page_widget.dart |
| 473 | 827 | 42.8% | clinicians new_a_p_p/member_details/pages/member_details_widget.dart |
| 469 | 835 | 43.8% | clinicians features/users/presentation/pages/user_details_page_widget.dart |
| 456 | 456 | 0.0% | clinicians new_a_p_p/new_messages/components/chat_home_component/chat_home_component_widget.dart |
| 447 | 806 | 44.5% | members onboarding/pages/referral_signup/referral_signup_widget.dart |
| 403 | 1175 | 65.7% | clinicians backend/api_requests/api_calls.dart |
| 402 | 1074 | 62.6% | members backend/api_requests/api_calls.dart |
| 397 | 397 | 0.0% | members main_pages/messages/components/chat_home_component/chat_home_component_widget.dart |
| 385 | 719 | 46.5% | members onboarding/pages/sso_signup_home/sso_signup_home_widget.dart |
| 374 | 988 | 62.1% | members onboarding/pages/login/login_widget.dart |
| 366 | 366 | 0.0% | clinicians new_a_p_p/appointment/appointment_call/appointment_call_widget.dart |
| 356 | 907 | 60.7% | clinicians new_a_p_p/components/responsive_side_menu/responsive_side_menu_widget.dart |
Priority within that list: release-critical flows first (auth, booking, payments, care flows, chat, documents), then per-feature domain and data layers, then presentation and provider logic.
Testing a legacy FlutterFlow page¶
The pages above are FlutterFlow-generated StatefulWidget + model pairs. What
they need, learned from the members login and SSO signup pages and the
clinicians members list and appointments pages:
- Pump through the app's
GoldenHarness.pumpApp. It wires Firebase, the auth manager,FFAppState, the localisation delegates and the silent network-image layer. Passsettle: falseand pump a fixed duration - these pages rarely reach a settled state. - Set
currentUserdirectly. Nothing listens to the auth stream in a widget test, and the clinicians side menu dereferencescurrentUserData!.permissions!unguarded, sosignInTestUseralone leaves the page unbuildable. - Stub the network at the seam the page actually uses. FlutterFlow
api_calls.dartentries go throughApiManager.instance.testClient(aMockClient); pages on the typed client overridetypedBffClinicalProviderwith a mockedTypedBffClinicalClient. - On clinicians, sign in properly or no request will arrive. Every
BFFClinicalGroupcall runs throughSessionInterceptor, which throwsToken not foundunless the auth manager holds a refresh token and an expiry - andFFApiInterceptor.makeApiCallturns a throwing interceptor into a synthetic 400 without ever calling the client. So a test that setscurrentUserby hand sees an enabled button, taps it, and captures nothing, with no error to explain why. UseGoldenHarness.signInTestUser(...), which sets all three, and passresetAuth: falsetopumpAppso it is not cleared. - Call
GoldenHarness.ignoreOverflowErrors()(clinicians). Ahem's glyphs are wider than the real font, so flex-less legacy rows overflow at paint time under test but not in the app. Everything else still fails the test. - A pending timer is usually a package's global state, not the page. The one
that blocked three of these widgets came from
visibility_detector, and no amount of reading the page would have found it: see below. Before hunting through a 900-line widget, read the pending-timer report the binding already prints.
When the load path misbehaves¶
The binding prints the duration, periodicity and creation stack of every pending timer immediately above the "A Timer is still pending" assertion. Read that first: it names the package and line that created it, which is faster and more reliable than inspecting the page.
Two package-level traps are already handled in
GoldenHarnessBase.baseGlobalSetUp (and mirrored in the members
GoldenHarness), so a page hitting either no longer needs anything per-test:
visibility_detectorschedules a static 500msTimerthe first time any detector paints. It surfaces in unhelpful ways: pumping past it does not clear it, because the callback repaints the detector, which schedules the next timer; and because the field is static only the first test in a process trips, so under--test-randomize-ordering-seed randomthe failure moves between tests and between runs.installVisibilityDetectorTestInterval()setsupdateIntervaltoDuration.zero, the package's own test path, which delivers callbacks from a post-frame callback and creates no timer.flutter_timezoneis the quieter one: unmocked, it does not settle within the frames a widget test pumps, so a page whose load path awaitsgetUserTimezone()- the clinicians appointments page does - never reaches its backend call. There is no error. The test sees no request at all, which reads like a provider override that did not take rather than a stalled dependency.installTimezoneMock()answers the channel.
A file that only trips sometimes is worse than no file, so fix the cause rather than shipping the test. Fix it in the harness when the cause is a package global: every future page test then inherits the fix.
Backend coverage¶
The backend stack is measured by .github/workflows/ci-backend.yml, not by the Flutter
gate above. functions and web-professionals each run vitest with the v8 coverage
provider and upload their lcov report to Datadog under their own flag - see
Datadog coverage reporting for the mechanics that apply to
every upload.
| Package | Flag | Line coverage (vitest) | Datadog gate |
|---|---|---|---|
functions |
backend-functions |
82.2% on develop | total_coverage_percentage >= 80 |
web-professionals |
backend-web-professionals |
73.10% (106/145) | none yet |
web-professionals baseline¶
Measured with pnpm run test --coverage in apps/perci-platform-backend/web-professionals
(vitest 4.1.8, 44 tests across 10 files, 12 source files in the report). Coverage counts
src/**/*.{ts,tsx} and excludes test files, *.d.ts and the test setup. The include glob is
deliberate on both counts: Vite compiles CSS and SVG imports to JS modules that the v8
provider would otherwise report as uncovered source, and without an include vitest measures
only the modules the tests actually imported, so untested files drop out of the denominator
and flatter the total. main.tsx, App.tsx and CalcomCalendar.tsx are untested and
datadog.ts is thinly covered; together they are the whole of the current gap.
Expect Datadog's figure for this flag to sit a little below the vitest one. vitest reports the
lcov LF: total (145 lines); lcov consumers commonly count the union of the DA: (line) and
BRDA: (branch) records instead, which here adds 6 lines carrying a branch record but no line
record - 4 in datadog.ts, 1 in main.tsx, 1 in CalcomAvailability.tsx - giving
106/151 = 70.20%. The hit count (106) is the same either way; only the denominator moves.
Quote whichever one you are actually looking at.
The two packages are not measured on the same basis, so do not read the rows as
like-for-like: functions/vitest.config.ts sets no coverage.include, so it counts only the
modules its tests imported. Aligning it will lower its reported number and move it under its
gate threshold, so it needs its own change and review.
web-professionals has no coverage gate yet. code-coverage.datadog.yml defines no
total_coverage_percentage gate for the backend-web-professionals flag, so a drop is
visible in Datadog but never blocks. Treat the baseline above as recorded, not protected. Add
a gate once the figure has settled, using the backend-functions entry as the template and
setting the floor just under the measured develop figure.
Awell care-flow lint¶
Awell care flows are the one production-critical part of the platform with no test gate:
they are authored in the Awell designer by non-engineers and published straight to the live
tenant with no commit, so PR CI structurally cannot catch a form or clinical-note action that
breaks. .github/workflows/awell-careflow-lint.yml runs
apps/perci-platform-backend/functions/src/scripts/lintAwellCareFlows.ts against the
production tenant nightly (06:00 UTC) and on demand (workflow_dispatch, run it before a
care-flow publish). It is deliberately not a pull-request check.
Four families of check, one report:
| Family | Check | Severity |
|---|---|---|
| Form definitions | mandatory question with no display rule after the first isLast exit screen |
failure |
rule references a question definition_id not in the form, or one later in the form |
failure | |
| empty title, duplicate question id | failure | |
| form title names a different cancer than its care flow | warning | |
| Clinical-note metadata | context key with surrounding whitespace (the backend reads the trimmed key, so the value is dropped) | failure |
type missing, or not an ActionContextType |
failure | |
boolean / priority / border_color / conclusion_status value the backend cannot parse |
failure | |
key the backend never reads (not in ActivityContextKey) |
warning | |
| API response schema | note sets no value for a field the action's schema requires (missing-required-value) |
failure |
note builds a value the action's schema rejects (schema-mismatch) |
failure | |
| Acuity scoring contract | a question the scorer reads is no longer asked, or an answer code it branches on is gone | failure |
| a question scored from a lookup table offers an option no score covers | failure | |
| two questions answer the same scored question, one question mixes two, or a question belongs to another domain | failure | |
| a question in an acuity form matches no scored question, or an "other, please specify" option has no free-text question after it | warning |
The API's own schemas decide the third family.
functions/src/scripts/checkNoteAgainstActionSchema.ts rebuilds each note into the action
getUnorderedActivities would send — stubbing only what the API supplies at runtime (activity
and form ids, pathway context, an article's WordPress copy, the practitioner profile,
appointment times) — and hands it to the real Zod schema in functions/src/schemas/actions.
Nothing parses those schemas at runtime (createApiRequestHandler validates requests only), so
an under-authored note never fails server-side: the builder substitutes "-", "" or
"unknown" and the member sees that. An article note with no resource names no WordPress
post, so its card opens nothing; a conclusion note with no resource sends an empty deeplink.
label is the deliberate exception — the schema types it non-nullable, but
AwellDashboardActivity.normalizedLabel in the members app reads the "-" back as absent and
promotes the title into the label slot, so an unset label is authored, not broken. Because a
schema change moves the goalposts for every published care flow, schemas/actions/** is on the
workflow's pull_request path filter.
The acuity family guards the scoring logic against form edits. Awell regenerates the
QuestionnaireResponse item linkIds on every publish, so the acuity scorer keys on the answer
codes instead: an option's value_string becomes the answer.valueCoding.code, and the code's
prefix names the question. An editor renaming an option, dropping one, or adding one to a
question scored from a lookup table changes what members score with no code change and no
error anywhere. functions/src/scripts/lintAcuityCareFlow.ts holds the live forms to
acuityScoringContract.ts, which lists every question the scorer reads, the codes it branches
on, and whether the question is scored by lookup (closed, so every option must be known) or
by count (open, so new options are safe).
Acuity care flows are recognised by content, not by id: a care flow offering at least three
distinct acuity questions is one, on whichever tenant and under whatever name. The trade-off is
that a care flow deleted outright is simply not found, so the run reports which domains it saw
and treats a domain nobody publishes as absent rather than broken. Set
AWELL_ACUITY_EXPECTED_DOMAINS=cancer,medical,social on the workflow once the forms are live in
production and a vanished care flow fails the run.
acuityScoringContract.ts is transcribed from common/acuity-scoring, not imported from it:
that module is on the PPL-3528 branch and not yet on develop. Two sources of truth is the known
cost of shipping this guard first; when the scorer lands, delete the literals and derive the
contract from ACUITY_DOMAIN_REGISTRY and acuityAnswerCodes instead.
Each finding says where to fix it. Form findings link straight to the form in Awell Studio
(/pathways/<care flow>/build/forms/<form definition id>). Clinical-note findings name the
track and the action (with its definition id) and link to the track builder; the orchestration
API does not expose an action's step definition id, and Studio's action route needs it, so the
track is the deepest reliable link. Clinical notes are read once per authored action: every
care-flow instance mints its own note from the same action_component.definition_id, and the
lint reads the newest one, which is the note from the latest published release.
No baseline: every defect keeps firing. The run exits non-zero and posts to #product-team
while any failure-severity finding stands, so it fails every night until the care flow is
fixed. A green run means the published care flows carry no known defect. Warnings
(title-cancer-mismatch, unknown-key) are listed in the job summary but never fail the run:
they are smells rather than broken behaviour, and unknown-key alone would drown the failures.
Findings are grouped by fingerprint, so one authoring slip is reported once however many forms
or notes reproduce it. A fingerprint is keyed on the care flow id, the form or question
definition_id, and the key/value, never on the form or note id: Awell mints a new form id on
every publish and every care-flow instance mints its own clinical note. To silence a finding,
fix the care flow in Awell Studio; there is no accept-and-move-on path. Run the sweep locally
before a publish:
cd apps/perci-platform-backend/functions
AWELL_API_KEY="$(gcloud secrets versions access latest --secret=AWELL_API_KEY --project=perci-platform-prod)" \
AWELL_API_ENDPOINT=https://api.uk.awellhealth.com/orchestration/m2m/graphql \
pnpm run awell:lint
Coverage is the latest published release only. forms(pathway_definition_id:) returns
the current version of each form, so a member mid-pathway on an older release can still hit a
defect the lint no longer sees, and fixing the latest version does not retroactively unblock
them. The report states this on every run; do not read a green run as a clean bill of health.
Retiring a care flow. Awell cannot unpublish a definition: once published it stays in
publishedPathwayDefinitions for good, and the lint keeps reporting defects in care flows the
product dropped long ago. To drop one from the sweep, the care-flow author sets retired in
its care-flow metadata in the Awell designer, then publishes that version and sets it live
(publishing records a version and does not touch any patient):
The lint skips those care flows and says how many it skipped in the job summary. Retire only a care flow nobody is running: members still mid-pathway keep hitting its forms, and a retired care flow is no longer checked. Metadata the parser cannot read leaves the care flow in the sweep, so a mistyped flag costs noise rather than coverage.