Skip to content

Testing & the CI/CD release gate

This page is the source of truth for how the Flutter apps (perci-platform-members and perci-platform-clinicians) are tested and what must be green before a change can be released. The goal is continuous delivery: a change merges and releases as soon as the gate is green, multiple times a day. The backend stack has its own gate and its own coverage flags - see Backend coverage at the end of this page.

Where QA testing happens (feature-branch first)

We work in feature branches: a long-lived epic branch is cut at the epic level, and each piece of implementation branches off it and PRs back into the epic branch (see Epic (feature) branches). Opening a PR with the preview label spawns a preview environment (members, clinicians, backend) for that branch.

Test the feature in its own branch's preview, before it merges. This is the in Test stage of the Delivery workflow; do not wait until the change is on develop / staging to verify it. By the time work reaches On Staging, that step is for regression against everything else, not first-pass feature testing.

The only exception: infrastructure that can't run in a preview

If a change depends on infrastructure that genuinely cannot be stood up in the feature branch's preview (for example a new queue, cron, external webhook, or environment-specific config), test it on staging after merge. This is the exception, not the norm, so call it out on the ticket to explain why it skipped feature-branch testing.

Test types

Flutter apps

Both apps use the same four layers, and the same tooling (patrol, golden_toolkit, mockito).

Layer Tool Lives in Runs in CI
Unit (domain/data) flutter_test test/features/**/domain, **/data every PR (blocking)
Widget flutter_test + ProviderScope overrides test/features/**/presentation every PR (blocking)
Golden golden_toolkit (@Tags(['golden'])) next to the widget, in goldens/ every PR (blocking)
E2E patrol (Chrome/web) patrol_test/ pre-release / on-label (separate)

Backend (functions)

Layer Tool Lives in Runs in CI
Unit / integration vitest functions/src/**/*.test.ts every PR (blocking)
Firestore security rules vitest + @firebase/rules-unit-testing + Firestore emulator functions/src/__tests__/firestore-rules/ every PR (blocking)

Run the rules suite locally:

cd apps/perci-platform-backend/functions
pnpm run test:rules

The suite covers member self-access, clinician self-access (notifications & calendar), cross-UID isolation (another member's or clinician's documents are denied), the users-public point-read-only policy, fully locked collections (plans, logs), and deny-by-default for unauthenticated clients. The rules and the config that drives the emulator both live in apps/perci-platform-backend/ (firestore.rules, firebase.json).

Mocking convention

  • Prefer hand-rolled fakes that implement the domain repository interface for unit/provider/widget tests - they are explicit, fast, and need no codegen.
  • Use mockito (@GenerateMocks + build_runner) for new tests where call verification or mocking a concrete SDK/generated client adds real value.
  • Providers are tested with a ProviderContainer (or ProviderScope) overriding the repository/datasource provider with a fake. For autoDispose providers, hold a listener before awaiting so the provider is not disposed mid-load.

The regression ratchet

A production regression closes with the test that would have caught it, at the lowest layer that can express the failure:

The bug is in… Test it with
A pure function, mapper, validator, or domain rule Unit
A request/response shape against a generated client or backend contract Unit against the client (a contract test), not a live call
Provider/notifier state, or a widget's reaction to state Widget (+ ProviderScope overrides)
Layout, sizing, or structure that silently collapsed Golden
Only the whole-flow interaction: navigation, auth, real cross-screen wiring E2E (patrol)

Pick the cheapest layer that actually fails on the bug. Reaching for E2E because it is quicker to write buys a slow test that catches less and flakes more; reaching for a unit test when the defect was the wiring between two screens proves nothing. The layer is a claim about where the defect lives, so state it on the ticket.

Backend regressions follow the same rule with vitest: a unit test over the handler or service, or a route-level test against the Zod schema, before anything that needs emulators.

The process around this is in Delivery. This section answers only the "which layer?" question.

Shared harness

packages/perci_platform_test_shared is the single source of truth for the fiddly, app-agnostic test setup: the silent network-image HTTP layer, the Firebase Analytics fake, the package_info / secure_storage / datadog channel mocks, golden_toolkit configuration, device presets and the Firebase core mocks. Each app keeps a thin test/.../golden_harness.dart that calls GoldenHarnessBase.baseGlobalSetUp() then wires its own Firebase init, auth manager and FFAppState. Patrol widget wrappers stay per-app (they embed each app's root widget).

The release gate (.github/workflows/ci-flutter.yml)

The two required PR status checks are Flutter checks passed (from ci-flutter.yml) and Backend checks passed (from ci-backend.yml). Both names are pinned in the branch rulesets — the workflow files can be renamed, the aggregate job names must not be.

A PR to develop or main must pass:

  1. Code generation - melos run build_runner (openapi + freezed + riverpod).
  2. Analyze (errors block) - flutter analyze --no-fatal-infos --no-fatal-warnings. Errors fail the build. Warnings/infos are reported but not yet fatal - they are a ratcheting backlog (see below). Flip to fatal-warnings once the count hits zero.
  3. Unit + widget tests - melos run test (excludes goldens), with coverage.
  4. Coverage threshold - total line coverage must be >= FLUTTER_MIN_COVERAGE (a repo variable). Ratchet this up toward 80%; never lower it.
  5. Goldens - a separate blocking job (see Goldens).

Patrol E2E runs in a separate workflow, not on every PR (see E2E).

Alongside the blocking checks, a non-blocking Code complexity job runs dart_code_linter over the app(s) in scope for the PR (members, clinicians, or both - the same scoping the other jobs use). The metric thresholds (cyclomatic complexity, maximum nesting level, number of parameters, source lines of code) live in each app's analysis_options.yaml. The tool is deliberately not registered as an analyzer plugin: a legacy plugin re-analyses the whole app inside the IDE's analysis server and splits the pub workspace into extra analysis contexts, which is what made in-IDE highlighting take 10-15 s (see "Analyzer performance" in running-locally.md). Metric violations in the files a PR touches appear as inline PR annotations (via scripts/frontend/dcl-github-annotations.dart — the tool's own GitHub reporter does not cover metrics); the full per-app reports are uploaded as the complexity-reports artifact, and when a PR's own files cross a threshold the job also maintains a sticky PR comment listing those findings - the detail a future blocking check would fail on. The job never blocks a merge. Run the same analysis locally with melos run complexity.

Coverage baseline & ratchet

Set the repo variable FLUTTER_MIN_COVERAGE to the current measured floor, then raise it as the backlog is burned down. The immediate purpose of the gate is non-regression (coverage may not drop); the long-term target is 80% on hand-written code, reached by ratcheting.

The floor follows the coverage actually present on the target branch. Raise it after each coverage-improving PR; never lower it below the branch's real coverage.

The merged report excludes generated Dart - *.g.dart, *.gr.dart, *.freezed.dart, *.mocks.dart and **/generated/** - so the figure the gate enforces and the figure Datadog reports measure the same thing. Both come from scripts/merge-flutter-coverage.sh, shared by ci-flutter.yml and coverage-develop.yml. On the PPL-3430 branch the exclusion removes 671 files and 13,155 lines.

Scope Line coverage
Merged, generated excluded (what the gate checks) 46.73% (38119/81569)
Merged, raw (for comparison with older figures) 42.04% (39870/94845)
perci-platform-members 46.75% (19543/41803)
perci-platform-clinicians 37.96% (14359/37829)
perci-platform-frontend-shared 39.23% (5968/15213)

PPL-3430 moved the merged figure from 25.97% to 46.73%. Set FLUTTER_MIN_COVERAGE to 46 once it is on develop.

perci-platform-frontend-shared alone has since moved from 39.23% to 80.94% (12600/15568), measured the same way (flutter test --coverage in the package, generated Dart excluded). The denominator grew from 15,213 to 15,568 because the new tests load libraries nothing loaded before - the effect the warning below describes. Re-measure the merged figure before moving FLUTTER_MIN_COVERAGE; the other two packages are unchanged by this work.

The denominator moves

flutter test --coverage only instruments libraries the tests actually load, so adding a test for a previously-untouched file adds its lines to the denominator as well as the numerator. A shallow smoke test over a large widget can therefore lower the reported total. Cover a chunk properly rather than merely touching it, and re-measure the whole suite - never a subset - before quoting a number.

The reported percentage is not coverage of the codebase

That mechanic has a bigger consequence than a wobbly number. A file no test loads appears in the report not at all - it is missing from the numerator and the denominator, so it is invisible rather than counted as zero. On the PPL-3430 branch 972 of 2,105 non-generated Dart files (46%) were absent entirely. The percentage is therefore coverage of the subset the suite happens to import, and it flatters: the files nobody has tested are exactly the ones excluded from the sum.

The fix is a library that imports everything, so the denominator is the whole package regardless of what the tests exercise. perci-platform-frontend-shared has one at test/coverage/all_libraries.dart; adding it moved that package from 37.27% over 200 files to 28.21% over 292 with the numerator unchanged. Lower, and true - and the shared figure in the table above is measured on that honest denominator, while the two app figures still are not.

Neither app can have one yet. lib/custom_code/actions/get_current_used_devices.dart in the shared package calls MediaStreamTrack.getSettings, which dart_webrtc provides on web but not on the Dart VM the test runner uses, and it is re-exported from the custom_code/actions barrel - so it is dragged into any complete compile of either app. Until it is made VM-safe or the barrel is split, the two app figures remain subset figures, and the merged number with them.

Datadog coverage reporting

Every suite uploads its merged lcov report to Datadog Code Coverage (EU site):

  • PRs - the datadog-ci coverage upload step in ci-flutter.yml and ci-backend.yml. On a pull_request event Datadog attributes the upload to GITHUB_HEAD_REF, so these only ever produce feature-branch coverage.
  • develop - .github/workflows/coverage-develop.yml re-runs the suites on push to develop and uploads, giving Datadog the default-branch baseline it diffs PR coverage against. Without it the develop branch view is empty.

The develop workflow deliberately runs tests only (no lint/build/threshold gate, no PR comment) and does not reuse the required check names, so it can never block a merge. Its env scaffolding mirrors the PR jobs; keep the two in sync or the baseline stops being comparable.

Every upload carries a flag - flutter, backend-functions or backend-web-professionals - set on the DataDog/coverage-upload-github-action step. Flags are what code-coverage.datadog.yml scopes carryforward to, and what the coverage gates will be scoped to. Most PRs touch one side of the repo only, so without them a backend-only PR reads as if all Flutter coverage had vanished. An unflagged upload merges into one undifferentiated pile that cannot be gated per stack. If a new package starts uploading, give it its own flag rather than reusing one of these.

code-coverage.datadog.yml (repo root) also maps paths to the Datadog services they already report under, so coverage lines up with APM, RUM and DORA, and ignores generated Dart output.

Datadog PR Gates

code-coverage.datadog.yml defines two PR Gates, both on total coverage. Patch coverage is measured in CI instead - see Patch coverage in CI.

Gate Scope Threshold Effect
total_coverage_percentage flag flutter 20 Fails the PR if Flutter total coverage drops below 20%
total_coverage_percentage flag backend-functions 80 Fails the PR if backend total coverage drops below 80%

threshold only accepts a number (0-100); there is no relative or auto value in the config file. A non-numeric threshold does not just disable that gate - it makes the whole file invalid, so services, ignore and carryforward stop applying too and the UI reports the config as invalid. The two totals above are floors set just under the current develop figures (backend 82.2%, Flutter 23.2%); raise them as coverage climbs, or express non-regression as a rule in the Datadog UI, which the file cannot do.

The totals are scoped by flag deliberately. Most PRs touch one side of the repo, and carryforward supplies the other side's figure from an ancestor commit. Scoping to the services instead would evaluate each one separately and fail PRs over packages the change never touched.

Scope multiplies evaluations, it does not filter them

services, codeowners and flags are all optional on a gate, and every scope listed is evaluated independently - "the gate does not combine coverage across services or code owners". Scope therefore decides how many times a gate runs, not whether it applies to the change in front of it. A scope covering code the PR never touched has no patch to measure, and:

Coverage is -1 when Datadog has no data - a no data sentinel, not 0% - and a gate reading -1 fails.

That is why there is no Datadog patch gate. Every scoping was tried and every one fails a PR that changes no coverable line, because there is genuinely nothing to measure:

Patch gate scope Result on a PR with no coverable changes
flags: ['*'] #1752 (a pubspec.yaml bump) - one red check
services: ['*'] #1896 (this file plus a doc) - four red checks, one per service
none #1898 - one red check

Scope changes how many times the gate goes red, never whether it applies. A gate that cannot be made blocking without making config-only PRs unmergeable is not a gate, so patch coverage moved to CI, where an empty diff can be skipped explicitly - see Patch coverage in CI. Do not re-add it here; the totals are the only thing this file gates.

Note that a gate is evaluated even when a PR's own CI skipped the test jobs and uploaded nothing - #1896's Flutter and backend aggregates passed off skips, and it was still gated. Path-filtering CI does not opt a PR out.

Datadog reports one status check per gate type, not per gate, so the two total gates share a single Total coverage percentage check. A check is only blocking once its context is added to the PR checks ruleset - a gate can fail while the PR stays mergeable until then, which is the case today for both. Gates also never apply retroactively: an open PR is only evaluated once a new commit is pushed to it. A stale Patch coverage percentage check may linger on PRs opened before that gate was removed; it is inert and disappears on the next push.

Gates can additionally be defined in the Datadog UI. Where the UI and this file both cover a scope, a PR must satisfy every threshold, so prefer keeping them here where they are reviewable alongside the code.

These sit alongside - not instead of - the in-CI gates: the vitest coverage.thresholds and the FLUTTER_MIN_COVERAGE step above still run, and still block. Both sets are floors, so keep them in step: the Datadog thresholds are evaluated on the merged report, the in-CI ones per package before upload.

Patch coverage in CI

scripts/patch-coverage.mjs replaces the Datadog patch gate, which could not be made blocking. It runs as a Patch coverage step in the Flutter and backend test jobs, after the lcov report is written and before it is uploaded, and:

  1. reads the lines the PR adds, from git diff --unified=0 HEAD^1 HEAD (HEAD^1 is the base branch tip on a pull_request merge-ref checkout, so the existing fetch-depth: 2 is enough history);
  2. keeps only those that appear as executable (DA:) lines in the lcov report;
  3. skips with a pass when that leaves nothing - the config-only case;
  4. otherwise fails below PATCH_MIN_COVERAGE (default 80).

Added lines in files with no lcov record - a YAML file, a Markdown doc, a package the suite does not instrument - are ignored rather than counted as missed. That is the whole difference from the Datadog gate, and the reason this one can be required.

It is warn-only until the PATCH_COVERAGE_ENFORCE repository variable is set to true; below-threshold PRs log a ::warning:: and pass. Flip the variable to make it blocking - no code change, and it is already inside the required Flutter checks passed / Backend checks passed aggregates, so nothing needs adding to the ruleset.

The script is plain dependency-free ESM rather than TypeScript because the Flutter job sets up no Node toolchain and runs it on the runner's stock node. Its unit tests are TypeScript, in scripts/__tests__/patch-coverage.test.ts, and run with the rest of the repo tooling via pnpm run test:release-tooling. Run it locally against a report you have already generated:

pnpm run patch-coverage -- --lcov coverage/lcov.info --label Flutter --base origin/develop

Coverage paths must be repo-relative. Datadog files a report by the paths in its SF: lines, so a package-relative report lands in a phantom top-level folder in the File Explorer. Flutter is handled by the merge step, which prefixes each package dir. Backend needs the same treatment: vitest runs inside the package and emits SF:src/..., so every backend job rewrites the paths into coverage/datadog/lcov.info before uploading (datadog-ci has no path-prefix flag), leaving the original lcov.info untouched. It is not only cosmetic - src/helpers/sleepSeconds.ts exists in both functions and web-professionals, so unprefixed the two reports collide on one phantom path. Any new package that starts uploading coverage needs the same rewrite.

Flaky test reporting

Datadog classifies a test as flaky when it sees both a passing and a failing status for the same commit. Datadog has no native Dart/Flutter tracer, so its Auto Test Retries and Early Flake Detection cannot run here - the JUnit XML upload in ci-flutter.yml is the only ingestion path, and the retry has to happen in the test run itself.

Each Flutter package that has tests therefore sets retry: 1 in its dart_test.yaml. A failing test gets a second attempt, so a one-off flake no longer reds the build and, more importantly, the pass/fail pair Datadog needs actually exists.

That alone is not enough. flutter test's JSON reporter collapses the attempts into a single entry - one testDone with result: success plus the error event from the attempt that failed - and junitreport:tojunit turns that into one <testcase> carrying a <failure>. Left as-is, a test that recovered on retry would be uploaded as a hard failure and never as a pass, so Datadog still could not classify it (and the data would be wrong).

scripts/frontend/flaky-junit-report.dart runs between the conversion and the upload and puts the missing attempt back. It reads the JSON report, which is the only source that distinguishes the two cases:

JSON result error events Meaning XML after the script
success >= 1 recovered on retry failing and passing testcase
failure one per attempt failed every attempt failing testcase, untouched
success none passed first time untouched

Datadog then sees a pass and a fail for the same test on the same commit and marks it flaky, which feeds the repo's existing auto-quarantine and auto-disable policies. Genuinely failing tests are left alone and still fail the build.

The E2E lane solves the same problem separately: Patrol runs with web-retries: 2 and reports a flaky-count (see The web lanes).

Analyze warning ratchet

melos analyze currently reports ~360 warnings/infos across the workspace, almost all pre-existing in legacy FlutterFlow code (perci_library_9rk85z) and a few in older test infra. There are no error-severity issues, so the error-only gate passes today. Burn the warning count down (a chunk is auto-fixable via dart fix --apply), then make warnings fatal in the gate.

Goldens

Golden tests run on every PR for both apps via flutter test --tags golden (the golden job in ci-flutter.yml). The job is blocking: a failing golden fails the required Flutter checks passed gate and prevents merge.

Cross-platform: render in boxes, not real fonts

Real fonts rasterise differently per OS (Windows DirectWrite, macOS CoreText, Linux FreeType), so golden PNGs made on one machine never match another - a mixed Win/Mac/Linux team plus Linux CI can't share real-font baselines. So our goldens do not load real fonts: the harness never calls loadAppFonts(), so Flutter's test environment renders all text in the Ahem font (every glyph a fixed square). Ahem output is identical on every platform, so text no longer moves a baseline between machines. (This is the same trick as Alchemist's "CI mode"; we do it directly rather than add the dependency.)

Ahem fixes text, not everything. Skia still rasterises shapes, borders and scaled images slightly differently across CPU architectures, so a baseline is not bit-identical between a dev machine and CI - see the rationale in each test/flutter_test_config.dart, which measured ~1.7% on Apple Silicon vs Linux x64 and set the comparator tolerance to 3% to clear it (8% in perci-platform-frontend-shared, whose goldens are more fill-heavy).

Consequence: goldens verify layout, sizing, colour and structure - not readable text (text shows as boxes). Text content is asserted by widget tests.

Linux CI is the authority, not your machine. A fill-heavy golden can sit just over the local tolerance and still be green on CI: as of 13 Aug 2026 twelve members goldens fail on macOS (plan_dashboard, mfa_security_page, mfa_setup_page, welcome_configuration, read_only_*, home_dashboard) while being green on the Linux runner, where the golden job gates every merge into develop. Before concluding your change broke a golden, run the suite with and without it - if the same set fails both ways, it is host drift. Regenerate baselines through the Flutter - Update Goldens workflow (which runs on CI) rather than locally, or you will bake macOS rasterisation into the baseline.

  • Regenerate through the Flutter - Update Goldens workflow, which runs on the Linux CI runner and opens a PR. flutter test <path> --tags golden --update-goldens locally is fine for iterating, but do not commit the result from a non-Linux machine: text is platform-independent, the rest is not, so you would trade your own drift for everyone else's.
  • Not golden-tested (non-deterministic regardless of fonts): live camera/video widgets (the old meeting_room / waiting_room goldens) and the animated WelcomePage were dropped; a couple of ultra-narrow scenarios are skip-ed where the wider Ahem glyphs tip a flex-less row into overflow (covered by their wider siblings + widget tests).

E2E (patrol)

Patrol E2E runs on four lanes, all selecting tests by tag:

Lane Workflow When Backend Gate?
Release candidate (web) e2e-web.yml every PR into main or a release/* branch the candidate's own preview backend fails its own run; not a required check
E2E authoring (web) e2e-web.yml any PR that edits patrol_test/, the shared harness, or the lane's CI definition real staging (the apps' committed endpoints) fails its own run; not a required check
Nightly (web) e2e-web-nightly.yml 02:30 UTC nightly + manual real staging advisory, deliberately not a required check
Native (Android) e2e-native-manual.yml every PR into main (automatic) + manual dispatch, any app, any branch real staging via Firebase Test Lab Native checks passed is a required check on main

Tag vocabulary

Lane selection is tag-driven. The vocabulary lives in one place — TestTags in packages/integration_test_shared — and raw-string vocabulary tags are rejected by a lint step in ci-flutter.yml (a typo'd raw tag would silently drop a test out of its lane). Zephyr case ids ('PPL-TNN') stay raw.

Tag Meaning Selected by
webSafe Honest-green, web-runnable, standalone (no seed step beyond what the lane's own action provides, no native device, non-destructive on shared staging) e2e-web.yml and e2e-web-nightly.yml (--tags webSafe)
native Needs a real Android device (Stripe CardField, WebRTC, permissions, file_picker) e2e-native-manual.yml — automatically on PRs into main, or via target / tags on manual dispatch
smoke / regression / desktop / tablet / mobile Suite descriptors; selectable ad hoc via the manual lanes' tags inputs —

Add webSafe to a case only once it is confirmed green and standalone; seed-dependent cases join when a seed pre-step lands. patrol bakes --tags into the generated bundle at build time, so each lane's build compiles exactly the matching cases.

The web lanes

Both apps share e2e-web.yml. It runs against Chrome (web) via Playwright and takes ~75 minutes, so it does not run on every PR — it has two lanes:

  • Release candidate — every pull request into main or a release/* branch, both apps unconditionally. A release ships both, so there is nothing to gain from filtering the matrix by which app's files changed.
  • E2E authoring — any pull request that edits the Patrol tests (apps/*/patrol_test/), their shared harness (packages/integration_test_shared/) or the lane's own CI definition, so a test change is proven by running it. Only the app(s) whose tests changed run; a shared or CI change runs both.

Every other pull request has no E2E lane: they gate on the unit, widget and contract suites, and the nightly run below is the net between releases.

A release candidate is tested against its own preview backend, not shared staging. The preview hosts are deterministic from the PR number, so this lane computes backend-pr-<N>.perci.dev itself and rewrites the app's environment.json — no handover from preview-pr.yml, which used to call this workflow (workflow_call) and tie every preview deploy to a 75-minute suite.

What it does still need from preview-pr.yml is when that backend is ready, so the Wait for the preview backend step polls that workflow's backend job for the PR's head sha. Only the job's conclusion separates the three outcomes: success is a fresh backend; skipped means the content-hash gate found the backend unchanged since the last push, so the preview host is already serving a content-identical build; failure means the host is still answering from the previous revision. The last case fails the E2E run rather than quietly falling back to staging, which would hand a release candidate a false green. That is also why the signal cannot be a commit-exact /_health probe: on a skipped deploy the preview backend legitimately reports an older sha. The job name is pinned in this workflow's PREVIEW_BACKEND_JOB_NAME env — renaming that job in preview-pr.yml without updating it turns the wait into a hard failure, deliberately.

The E2E-authoring lane has no preview backend to wait for and tests the staging endpoints committed in each app's environment.json.

The nightly lane (e2e-web-nightly.yml) runs the same webSafe subset for both apps against real staging, on the 02:30 UTC cron or on demand. It is deliberately not wired to pull requests: a PR into main or release/* already runs that subset through the lane above, and that lane is the one that publishes the Zephyr cycle. Triggering both ran every case twice per release push. The nightly lane keeps one thing the release-candidate lane cannot have — it uploads its JUnit report to Datadog Test Optimization, which needs a Datadog key that e2e-web.yml is deliberately never handed, because that lane builds and runs PR-authored code.

SSO sign-up: the identity-provider sidecar

The member SSO sign-up case (sso_sign_up_test.dart, Zephyr T4) is the one web case that cannot stay inside the Flutter page: the real flow leaves the app for the identity provider, and the Descope exchange code the provider hands back is single use and expires within minutes, long before a web bundle finishes building. So the identity leg runs in a sidecar the test calls at run time. patrol_test/tools/sso/serve_sso_mint.mjs enrols a throwaway account on the team's test identity provider (Authentik at auth.perci.dev, whose enrolment flow is public, so no credentials or secrets are involved), performs the real login and consent headlessly, captures the code from the redirect without loading it, and returns it with the identity the app should prefill. The test then enters at /sso-callback?code=…, the exact route the browser redirect uses, and drives the sign-up forms to the dashboard, asserting the identity the provider supplied is what the app prefilled.

.github/actions/patrol-web-test starts the sidecar for the members app and passes --dart-define=SSO_MINT_URL=…; without that define the case is skipped, never failed. Locally, zsh patrol_test/tools/run_sso_patrol.sh does the same. The sidecar mints against the backend the app's environment.json points at, so a release candidate exchanges codes its own preview backend produced. Each run leaves one sso-e2e-<epoch> Authentik user and one staging member behind, the same footprint as the voucher sign-up case. Details, overrides and hygiene notes: apps/perci-platform-members/patrol_test/tools/sso/README.md.

Zephyr: the automated release cycle

QA plans and records release testing in Zephyr Scale, so the automated coverage has to land there too rather than living only in a CI log. On a release/* or hotfix/* PR into main, the zephyr job in e2e-web.yml publishes the web lane's results as Zephyr test cycles, one per app, role and device, with each case marked Pass or Fail:

Release <x.y.z> - <Clinician|Member> App - <Role, if applicable> - <Desktop|Mobile> - Automated

That mirrors the manual cycles (Release 1.1.23 - Clinician App - Nurse Admin - Manual): app first, then role where one applies, then device, with the execution mode last.

The app comes from the artefact the report arrived in (members-test-logs / clinicians-test-logs), since a spec title does not carry it. The device comes from the lane: the Chrome lane is a single 1920x1080 viewport so it publishes Desktop, and Firebase Test Lab runs real devices so it will publish Mobile (PPL-3444). The existing desktop/tablet tags are not used for this, because they describe intent rather than the viewport that actually ran.

Every cycle is attached to the release's Zephyr Test Plan, named Release <x.y.z> to match the plans QA already keeps (Release 1.1.20 upward). The plan is created if the release does not have one yet, so the automated cycles hang off the same plan as the manual ones.

Worth knowing if you touch that code: the link is written from the plan side (POST /testplans/{planKey}/links/testcycles) and the field is testCycleIdOrKey, even though the payloads report the relationship as testCycleId and a cycle reads its own links back from GET /testcycles/{key}/links. The published API docs are unreachable, so both facts came from the live API; the obvious reading of the payloads gives the wrong side and the wrong field name.

If the plan cannot be created or linked, the cycles and executions publish anyway and the run warns. A plan groups the record, whereas the executions are the record.

The role comes from the test's own TestTags.role* tag, and this is the part worth understanding. A Zephyr case is labelled for every role a manual tester covers with it: PPL-T47 carries Nurse, Admin_Nurse and AHP. An automated test signs in as one. Grouping on the case's labels would therefore record a Pass in the AHP cycle for a run that never signed in as AHP, so instead each test declares the role it actually runs as and only ever lands in that role's cycle. A test that genuinely signs in as two roles carries both tags and appears in both.

Tag the role the test really exercises:

patrolTest(
  '[PPL-T37] Appointments: clinician views their appointment list',
  tags: [TestTags.regression, TestTags.webSafe, 'PPL-T37', TestTags.roleNurse],

For a test parameterised over CliniciansRoles.roles, use TestTags.roleTagFor(role) so each generated test carries only its own role. A case that signs in as nobody (the member app, or PPL-T10 which only asserts the pre-auth sign-in UI) carries no role tag and lands in the app's roleless cycle.

Note that the role names are not covered by the raw-string tag lint in ci-flutter.yml. Unlike the lane vocabulary, 'Nurse' is also a legitimate credentialRegistry['Nurse'] key, so a single-line grep cannot tell a stray raw tag from a login call.

The link between a Patrol test and its Zephyr case is the case id in the test name:

patrolTest(
  '[PPL-T25] Member details Message button opens the conversation',
  tags: [TestTags.regression, TestTags.desktop, 'PPL-T25'],

patrol's web runner uses the Dart test name as the Playwright spec title, so scripts/zephyr/publish-e2e-results.ts reads the case ids straight out of the playwright-report/results.json the Patrol job already uploads. Both authored shapes are understood — [PPL-T196] [PPL-T187] and the combined [PPL-T47 T48 T49]. A case with no id in its title is invisible to Zephyr, which is the main thing to get right when adding a test.

Each execution also carries:

  • The measured run time, in Zephyr's "Actual Time" field, taken from the attempt that decided the outcome. For a case that passed on retry that is the successful attempt, so the figure is the test's own run time rather than the sum of its failed attempts.
  • The failure message in the execution comment, with Playwright's ANSI colouring stripped and the text escaped and truncated so a long assertion diff cannot break the rendering.
  • An evidence pointer, on every execution. Pass or fail, the comment names the CI artefact (clinicians-test-logs), links the run, and says to open playwright-report/index.html. A failure additionally quotes its video path. Reviewing a failure usually means comparing it against a passing run, so an execution with no pointer at all just forces a hunt through the Actions UI.

Playwright records attachment paths as absolute paths on the machine that ran the suite, so the video path is re-anchored onto the downloaded artefact by its test-results/ segment, and is only quoted once the file has actually been found.

Evidence is referenced, not attached. Zephyr Scale's v2 API rejects POST /testexecutions/{id}/attachments with 405 Method Not Allowed, so nothing can be uploaded from CI. The reference is more useful anyway: the artefact holds the HTML report, every video and the full log, not just one file.

Because those references are the release's evidence and QA reviews the cycle after the release ships, a release or hotfix candidate keeps its artefacts for 90 days instead of the ordinary 7 (see retention-days in e2e-web.yml). At 7 days every citation in Zephyr would rot within the week.

Playwright traces are not produced today: patrol's playwright.config.ts sets video but never trace, and neither its env vars nor patrol_cli expose a way to turn it on. Enabling them needs a change upstream in patrol.

Behaviour worth knowing:

  • A case that passed only on retry is recorded as Pass with a flaky note, matching the lane's own verdict; a case that failed every attempt is a Fail.
  • A combined case ([PPL-T47 T48 T49]) attaches the same video to each of its executions, since each case is a separate row and each should carry its own evidence.
  • Skipped cases get no execution at all — the run says nothing about them, and a row implying otherwise is worse than no row.
  • Re-running the workflow reuses the existing cycle for that version instead of creating a duplicate.
  • Publishing is not a gate. It never fails the E2E verdict, and a missing ZEPHYR_API_TOKEN repository secret only logs a warning.
  • A case id present in the suite but unknown to Zephyr is reported in the job summary rather than dropped silently, so the two inventories can be reconciled.

The version comes from the PR head's apps/perci-platform-members/pubspec.yaml, not the branch name, so hotfix branches (which carry no version in the branch name) name their cycle correctly.

The cycle is also linked to the Jira release version, so QA can filter cycles by release. The lookup is by exact name and tries Platform <x.y.z> before a bare <x.y.z>, because PPL contains both conventions plus unrelated Member App x.y.z and Clinical App x.y.z versions that would make a loose semver match pick the wrong release. It needs the JIRA_BASE_URL, JIRA_USER_EMAIL and JIRA_API_TOKEN secrets; without them, or when no version matches, the cycle is still published, just unlinked.

Cases covered here should carry the automated label in Zephyr, which is what keeps /zephyr-release-cycles from also pulling them into the manual cycles.

The native lane publishes too

e2e-native-manual.yml publishes into the same cycles, as Mobile. gcloud firebase test android run only prints a leg-level pass or fail, but FTL writes a test_result_*.xml per device and shard into its results bucket, with a <testcase> per Dart test. Each leg pins --results-dir so the path is predictable, recovers the bucket name from FTL's own output, and gsutil cps the XML into its artefact.

Patrol names each native test <filename> [PPL-T93] <title>, because its JUnit runner is parameterised over the Dart test names, so the same key extraction works unchanged.

Two things differ from the web lane:

  • JUnit XML carries no tags, so the app and role cannot come from the results. Each leg runs pnpm run zephyr:case-meta against the commit under test and ships the resulting case-meta.json beside the XML; the publisher joins on it. That is also why the map is generated in the test job rather than the publish job, which checks out the base branch.
  • The device is derived from the lane, not a flag. FTL runs real devices, so its results are always Mobile. Deriving it per result means a single run covering both lanes cannot mislabel the native half, which one global --device would.

The evidence pointer differs too: an FTL artefact holds JUnit XML, logcat and the device video, so a native comment points at the Firebase Test Lab results rather than a playwright-report that was never there.

The XML shape is FTL's standard instrumentation output and is parsed tolerantly (both <testsuites><testsuite> and a bare <testsuite>, failure or error, skipped), but it has not yet been confirmed against a live run.

Dry-run the mapping locally against a downloaded report:

pnpm run zephyr:publish-e2e -- --reports ./artifacts --version 1.1.30 --dry-run

The native lane

Some flows cannot run in headless Chrome at all — Stripe's CardField is an Android platform view, WebRTC and GetStream need real media permissions, file_picker needs a real file system. Those cases are tagged native and run on real Android devices in Firebase Test Lab via e2e-native-manual.yml.

Every PR into main — so every release and hotfix candidate — runs the native matrix automatically, gated by the required Native checks passed check. The matrix is over (app, device, selector), not just app, because the cases disagree about the device they need:

Leg Device Selects Seeds
members purchase MediumPhone.arm portrait --tags native (PPL-T100, Stripe CardField) seed_screening_purchase.sh, re-run per run because the pathway is consumed
clinicians video / in-call MediumTablet.arm landscape --tags "native && !PPL-T140" seed_join_video_appointment.mjs
clinicians file picker MediumTablet.arm landscape --target document_upload_native_ppl_t140_test.dart none; pushes the fixture PDF via --other-files

The clinicians legs need a tablet in landscape because the member-details page is an unconditional two-column layout with no mobile breakpoint — on a phone viewport the right panel is crushed and its tap targets go off-screen.

Seeding now happens inside the workflow, so no manual pre-step is needed. Manual dispatch still takes either app, any branch, a single target or a tags selector, with optional FTL sharding for parallelism.

The call-member cluster

These four cases look alike and are routinely lumped together, but they belong on different lanes. Only one of them places a real call:

Case Places a real call? Lane Needs
PPL-T91/T96 call_member_connect yes web only the auto-answer Twilio peer
PPL-T92 cancel modal no, confirm is asserted but never tapped web a member with a phone
PPL-T93 mic denied no native only the FTL harness withholding RECORD_AUDIO
PPL-T97 call_member_tier1 no web a member with a phone

PPL-T91/T96 cannot run on the native lane at all: the app never registers an Android Telecom PhoneAccount, so the twilio_voice plugin rejects makeCall with "No registered phone account". The clinician calling path is web-primary, and the case is honest-green on the web lane because --use-fake-device-for-media-stream lets the Twilio Voice JS SDK place the call and it genuinely connects to the auto-answer peer. It is deliberately not webSafe — it places a real, billed PSTN call, so it does not belong on a per-PR lane. See PPL-3379.

PPL-T93 is the mirror image: native-only, because the androidTest harness grants RECORD_AUDIO per Dart test name and deliberately withholds it for this one case (MainActivityTest.java), which is what makes the denied path reachable. It needs no Twilio number and places no call, so it runs on the automatic native lane.

Every user story should land with an automated E2E that walks its flow, so the patrol suite grows into a full regression net. That suite is the release regression gate: because it runs on every PR into main or a release/* branch, a green run is what lets us ship confident there are no regressions, without a manual re-test pass. Adding the E2E is part of the Definition of Done.

The shared harness lives in patrol_test/ (patrol_setup.dart, patrol_widget_wrapper.dart, helpers/clinician_session.dart, pages/). patrol test discovers patrol_test/ by default. Clinician flows covered: sign-in (+ forgot-password), sign-out, members-list search, members-list filter, member details (contact + demographic), member medical record (documents + screening sections), appointments (tabs), messages, payments, and the Learn hub (open + search). Section/tab flows are permission-gated and skip cleanly when a role lacks access. These are authored + analyze-verified; runtime execution is the patrol CI job.

Autofix on failure

When a Patrol run fails, an agentic follow-up (.github/workflows/e2e-web-autofix.yml) investigates instead of waiting for a human's slow edit-and-rerun loop. It decides whether the failure is a flaky test (e.g. a waitUntilVisible hitting a refresh animation) or a real functionality break — probing live UI state with marionette and re-running only the failing target as the authoritative gate — then opens a draft PR with the fix (test-side for a flake; a minimal product fix plus a report for a real break). See Patrol autofix for the design.

Running locally

melos bootstrap            # resolve deps (first time / after pubspec changes)
melos run build_runner     # generate openapi + freezed + riverpod
melos run analyze          # static analysis
melos run test             # unit + widget (excludes goldens)

# single app, with coverage
cd apps/perci-platform-members && fvm flutter test --coverage --exclude-tags golden
# goldens for one file
fvm flutter test <path> --tags golden            # compare
fvm flutter test <path> --tags golden --update-goldens   # regenerate

Python: the SFTP referral importer

The only Python in the repo is the referral-import Cloud Function at infrastructure/sftp-referrals/functions/process-referral/, which parses partner referral PDFs and posts them to the EBP Referral API. It has its own harness because it shares nothing with the Flutter or Node stacks.

Layer Tool Lives in Runs in CI
Unit pytest test_*.py, beside the code every PR (blocking)
cd infrastructure/sftp-referrals/functions/process-referral
pip install -r requirements-dev.txt
./run_tests.sh          # what CI runs — do not use bare `pytest`, see below

Tests sit next to the module rather than in a tests/ subdirectory, so from parsers.base import ... resolves without packaging the function — pytest prepends the test file's directory to sys.path. requirements-dev.txt pulls in the full runtime requirements on purpose: the parser imports pdfplumber at module scope, and installing the real dependency set means CI catches an import that only resolves on someone's laptop. That is the failure mode behind the exact pin on rapidocr-onnxruntime (PPL-2200).

Individual files also run standalone (python3 test_slack_notify.py) via a __main__ block. That is a convenience, not the contract — CI runs run_tests.sh.

One interpreter per test file

run_tests.sh invokes pytest once per test_*.py, and bare pytest over the directory does not work. Every test file fakes the modules main.py imports (requests, google.cloud.storage, functions_framework, parsers) by installing stubs into sys.modules before import main.

In one process main is imported once, and it binds whichever stubs were installed at that moment. The first file pytest collects wins; every later file's stubs land in sys.modules too late to reach main. The symptom is a file whose assertions see an empty capture list, because main is posting into another file's fake — which reads like a bug in the code under test and is not one. main.py also builds a module-level storage_client at import time, which no amount of later rebinding can redirect.

Isolation is what these tests were written to expect, so the harness gives it to them. The proper fix is a shared conftest.py installing one set of stubs for the whole suite (PPL-3418); when that lands, this script can go.

If you add a test file that stubs a module main imports, keep the stub module-level and do not rely on another file's fakes being visible.

Fixtures: never use real referral data

Referral PDFs are member clinical records. Names, policy numbers, claim numbers and email addresses in tests are always synthetic — do not paste values from production, from a bucket object name, or from a log line into a test file. Committed test data lives in git history forever.

Parser logic is therefore factored into pure string helpers that need no PDF at all, which is how the majority of cases are covered. If a test genuinely needs a PDF, generate a synthetic one at test time matching the form's table layout; do not commit a real referral, redacted or otherwise.

The gate

The job is SFTP referrals checks in ci-backend.yml, gated on a sftp_referrals paths filter and wired into the Backend checks passed aggregator. It reuses that existing required check rather than adding a new one, so no branch-ruleset change is needed for it to block a merge.

Before this existed, a PR touching only infrastructure/sftp-referrals/** saw Backend checks passed and Flutter checks passed both go green in nine seconds off skipped jobs, with no test executed. If you add another stack outside apps/**, give it a filter and add it to an aggregator's needs — a required check that passes because everything under it was skipped is worse than no check, because it reads as assurance.

Coverage backlog (path to "all functionality covered")

Comprehensive coverage is delivered by ratcheting FLUTTER_MIN_COVERAGE, not in a single pass. Work is chunked by page or feature, one chunk per PR, each measured against a full-suite run.

Regenerate this ranking from a merged report at any time:

awk -F: '/^SF:/{f=substr($0,4);h=0;n=0} /^DA:/{split(substr($0,4),a,",");n++;if(a[2]+0>0)h++} /^end_of_record/{if(n>0)printf "%6d %6d %5.1f%% %s\n", n-h, n, 100*h/n, f}' coverage/lcov.info | sort -rn | head -40

Largest remaining debts

Measured 13 Aug 2026, after the PPL-3430 chunks. Uncovered lines first, since those are lines already in the denominator - covering them raises the total without moving the goalposts.

Uncovered Total Cov File
956 956 0.0% clinicians new_a_p_p/calendar/pages/calendar_home/calendar_home_widget.dart
591 856 31.0% shared components/address_field_library/address_field_library_widget.dart
511 948 46.1% clinicians new_a_p_p/your_details/personal_details/personal_details_widget.dart
486 905 46.3% clinicians new_a_p_p/appointment/appointments_page/appointments_page_widget.dart
473 827 42.8% clinicians new_a_p_p/member_details/pages/member_details_widget.dart
469 835 43.8% clinicians features/users/presentation/pages/user_details_page_widget.dart
456 456 0.0% clinicians new_a_p_p/new_messages/components/chat_home_component/chat_home_component_widget.dart
447 806 44.5% members onboarding/pages/referral_signup/referral_signup_widget.dart
403 1175 65.7% clinicians backend/api_requests/api_calls.dart
402 1074 62.6% members backend/api_requests/api_calls.dart
397 397 0.0% members main_pages/messages/components/chat_home_component/chat_home_component_widget.dart
385 719 46.5% members onboarding/pages/sso_signup_home/sso_signup_home_widget.dart
374 988 62.1% members onboarding/pages/login/login_widget.dart
366 366 0.0% clinicians new_a_p_p/appointment/appointment_call/appointment_call_widget.dart
356 907 60.7% clinicians new_a_p_p/components/responsive_side_menu/responsive_side_menu_widget.dart

Priority within that list: release-critical flows first (auth, booking, payments, care flows, chat, documents), then per-feature domain and data layers, then presentation and provider logic.

Testing a legacy FlutterFlow page

The pages above are FlutterFlow-generated StatefulWidget + model pairs. What they need, learned from the members login and SSO signup pages and the clinicians members list and appointments pages:

  • Pump through the app's GoldenHarness.pumpApp. It wires Firebase, the auth manager, FFAppState, the localisation delegates and the silent network-image layer. Pass settle: false and pump a fixed duration - these pages rarely reach a settled state.
  • Set currentUser directly. Nothing listens to the auth stream in a widget test, and the clinicians side menu dereferences currentUserData!.permissions! unguarded, so signInTestUser alone leaves the page unbuildable.
  • Stub the network at the seam the page actually uses. FlutterFlow api_calls.dart entries go through ApiManager.instance.testClient (a MockClient); pages on the typed client override typedBffClinicalProvider with a mocked TypedBffClinicalClient.
  • On clinicians, sign in properly or no request will arrive. Every BFFClinicalGroup call runs through SessionInterceptor, which throws Token not found unless the auth manager holds a refresh token and an expiry - and FFApiInterceptor.makeApiCall turns a throwing interceptor into a synthetic 400 without ever calling the client. So a test that sets currentUser by hand sees an enabled button, taps it, and captures nothing, with no error to explain why. Use GoldenHarness.signInTestUser(...), which sets all three, and pass resetAuth: false to pumpApp so it is not cleared.
  • Call GoldenHarness.ignoreOverflowErrors() (clinicians). Ahem's glyphs are wider than the real font, so flex-less legacy rows overflow at paint time under test but not in the app. Everything else still fails the test.
  • A pending timer is usually a package's global state, not the page. The one that blocked three of these widgets came from visibility_detector, and no amount of reading the page would have found it: see below. Before hunting through a 900-line widget, read the pending-timer report the binding already prints.

When the load path misbehaves

The binding prints the duration, periodicity and creation stack of every pending timer immediately above the "A Timer is still pending" assertion. Read that first: it names the package and line that created it, which is faster and more reliable than inspecting the page.

Two package-level traps are already handled in GoldenHarnessBase.baseGlobalSetUp (and mirrored in the members GoldenHarness), so a page hitting either no longer needs anything per-test:

  • visibility_detector schedules a static 500ms Timer the first time any detector paints. It surfaces in unhelpful ways: pumping past it does not clear it, because the callback repaints the detector, which schedules the next timer; and because the field is static only the first test in a process trips, so under --test-randomize-ordering-seed random the failure moves between tests and between runs. installVisibilityDetectorTestInterval() sets updateInterval to Duration.zero, the package's own test path, which delivers callbacks from a post-frame callback and creates no timer.
  • flutter_timezone is the quieter one: unmocked, it does not settle within the frames a widget test pumps, so a page whose load path awaits getUserTimezone() - the clinicians appointments page does - never reaches its backend call. There is no error. The test sees no request at all, which reads like a provider override that did not take rather than a stalled dependency. installTimezoneMock() answers the channel.

A file that only trips sometimes is worse than no file, so fix the cause rather than shipping the test. Fix it in the harness when the cause is a package global: every future page test then inherits the fix.

Backend coverage

The backend stack is measured by .github/workflows/ci-backend.yml, not by the Flutter gate above. functions and web-professionals each run vitest with the v8 coverage provider and upload their lcov report to Datadog under their own flag - see Datadog coverage reporting for the mechanics that apply to every upload.

Package Flag Line coverage (vitest) Datadog gate
functions backend-functions 82.2% on develop total_coverage_percentage >= 80
web-professionals backend-web-professionals 73.10% (106/145) none yet

web-professionals baseline

Measured with pnpm run test --coverage in apps/perci-platform-backend/web-professionals (vitest 4.1.8, 44 tests across 10 files, 12 source files in the report). Coverage counts src/**/*.{ts,tsx} and excludes test files, *.d.ts and the test setup. The include glob is deliberate on both counts: Vite compiles CSS and SVG imports to JS modules that the v8 provider would otherwise report as uncovered source, and without an include vitest measures only the modules the tests actually imported, so untested files drop out of the denominator and flatter the total. main.tsx, App.tsx and CalcomCalendar.tsx are untested and datadog.ts is thinly covered; together they are the whole of the current gap.

Expect Datadog's figure for this flag to sit a little below the vitest one. vitest reports the lcov LF: total (145 lines); lcov consumers commonly count the union of the DA: (line) and BRDA: (branch) records instead, which here adds 6 lines carrying a branch record but no line record - 4 in datadog.ts, 1 in main.tsx, 1 in CalcomAvailability.tsx - giving 106/151 = 70.20%. The hit count (106) is the same either way; only the denominator moves. Quote whichever one you are actually looking at.

The two packages are not measured on the same basis, so do not read the rows as like-for-like: functions/vitest.config.ts sets no coverage.include, so it counts only the modules its tests imported. Aligning it will lower its reported number and move it under its gate threshold, so it needs its own change and review.

web-professionals has no coverage gate yet. code-coverage.datadog.yml defines no total_coverage_percentage gate for the backend-web-professionals flag, so a drop is visible in Datadog but never blocks. Treat the baseline above as recorded, not protected. Add a gate once the figure has settled, using the backend-functions entry as the template and setting the floor just under the measured develop figure.

Awell care-flow lint

Awell care flows are the one production-critical part of the platform with no test gate: they are authored in the Awell designer by non-engineers and published straight to the live tenant with no commit, so PR CI structurally cannot catch a form or clinical-note action that breaks. .github/workflows/awell-careflow-lint.yml runs apps/perci-platform-backend/functions/src/scripts/lintAwellCareFlows.ts against the production tenant nightly (06:00 UTC) and on demand (workflow_dispatch, run it before a care-flow publish). It is deliberately not a pull-request check.

Four families of check, one report:

Family Check Severity
Form definitions mandatory question with no display rule after the first isLast exit screen failure
rule references a question definition_id not in the form, or one later in the form failure
empty title, duplicate question id failure
form title names a different cancer than its care flow warning
Clinical-note metadata context key with surrounding whitespace (the backend reads the trimmed key, so the value is dropped) failure
type missing, or not an ActionContextType failure
boolean / priority / border_color / conclusion_status value the backend cannot parse failure
key the backend never reads (not in ActivityContextKey) warning
API response schema note sets no value for a field the action's schema requires (missing-required-value) failure
note builds a value the action's schema rejects (schema-mismatch) failure
Acuity scoring contract a question the scorer reads is no longer asked, or an answer code it branches on is gone failure
a question scored from a lookup table offers an option no score covers failure
two questions answer the same scored question, one question mixes two, or a question belongs to another domain failure
a question in an acuity form matches no scored question, or an "other, please specify" option has no free-text question after it warning

The API's own schemas decide the third family. functions/src/scripts/checkNoteAgainstActionSchema.ts rebuilds each note into the action getUnorderedActivities would send — stubbing only what the API supplies at runtime (activity and form ids, pathway context, an article's WordPress copy, the practitioner profile, appointment times) — and hands it to the real Zod schema in functions/src/schemas/actions. Nothing parses those schemas at runtime (createApiRequestHandler validates requests only), so an under-authored note never fails server-side: the builder substitutes "-", "" or "unknown" and the member sees that. An article note with no resource names no WordPress post, so its card opens nothing; a conclusion note with no resource sends an empty deeplink. label is the deliberate exception — the schema types it non-nullable, but AwellDashboardActivity.normalizedLabel in the members app reads the "-" back as absent and promotes the title into the label slot, so an unset label is authored, not broken. Because a schema change moves the goalposts for every published care flow, schemas/actions/** is on the workflow's pull_request path filter.

The acuity family guards the scoring logic against form edits. Awell regenerates the QuestionnaireResponse item linkIds on every publish, so the acuity scorer keys on the answer codes instead: an option's value_string becomes the answer.valueCoding.code, and the code's prefix names the question. An editor renaming an option, dropping one, or adding one to a question scored from a lookup table changes what members score with no code change and no error anywhere. functions/src/scripts/lintAcuityCareFlow.ts holds the live forms to acuityScoringContract.ts, which lists every question the scorer reads, the codes it branches on, and whether the question is scored by lookup (closed, so every option must be known) or by count (open, so new options are safe).

Acuity care flows are recognised by content, not by id: a care flow offering at least three distinct acuity questions is one, on whichever tenant and under whatever name. The trade-off is that a care flow deleted outright is simply not found, so the run reports which domains it saw and treats a domain nobody publishes as absent rather than broken. Set AWELL_ACUITY_EXPECTED_DOMAINS=cancer,medical,social on the workflow once the forms are live in production and a vanished care flow fails the run.

acuityScoringContract.ts is transcribed from common/acuity-scoring, not imported from it: that module is on the PPL-3528 branch and not yet on develop. Two sources of truth is the known cost of shipping this guard first; when the scorer lands, delete the literals and derive the contract from ACUITY_DOMAIN_REGISTRY and acuityAnswerCodes instead.

Each finding says where to fix it. Form findings link straight to the form in Awell Studio (/pathways/<care flow>/build/forms/<form definition id>). Clinical-note findings name the track and the action (with its definition id) and link to the track builder; the orchestration API does not expose an action's step definition id, and Studio's action route needs it, so the track is the deepest reliable link. Clinical notes are read once per authored action: every care-flow instance mints its own note from the same action_component.definition_id, and the lint reads the newest one, which is the note from the latest published release.

No baseline: every defect keeps firing. The run exits non-zero and posts to #product-team while any failure-severity finding stands, so it fails every night until the care flow is fixed. A green run means the published care flows carry no known defect. Warnings (title-cancer-mismatch, unknown-key) are listed in the job summary but never fail the run: they are smells rather than broken behaviour, and unknown-key alone would drown the failures.

Findings are grouped by fingerprint, so one authoring slip is reported once however many forms or notes reproduce it. A fingerprint is keyed on the care flow id, the form or question definition_id, and the key/value, never on the form or note id: Awell mints a new form id on every publish and every care-flow instance mints its own clinical note. To silence a finding, fix the care flow in Awell Studio; there is no accept-and-move-on path. Run the sweep locally before a publish:

cd apps/perci-platform-backend/functions
AWELL_API_KEY="$(gcloud secrets versions access latest --secret=AWELL_API_KEY --project=perci-platform-prod)" \
AWELL_API_ENDPOINT=https://api.uk.awellhealth.com/orchestration/m2m/graphql \
pnpm run awell:lint

Coverage is the latest published release only. forms(pathway_definition_id:) returns the current version of each form, so a member mid-pathway on an older release can still hit a defect the lint no longer sees, and fixing the latest version does not retroactively unblock them. The report states this on every run; do not read a green run as a clean bill of health.

Retiring a care flow. Awell cannot unpublish a definition: once published it stays in publishedPathwayDefinitions for good, and the lint keeps reporting defects in care flows the product dropped long ago. To drop one from the sweep, the care-flow author sets retired in its care-flow metadata in the Awell designer, then publishes that version and sets it live (publishing records a version and does not touch any patient):

{ "retired": true }

The lint skips those care flows and says how many it skipped in the job summary. Retire only a care flow nobody is running: members still mid-pathway keep hitting its forms, and a retired care flow is no longer checked. Metadata the parser cannot read leaves the care flow in the sweep, so a mistyped flag costs noise rather than coverage.