Skip to content

Datadog error triage

Error Tracking is only useful if env:prod means "real users in production". This page records which environments are allowed to report, and where each gate lives.

Alerting tiers

There is currently one tier, and it runs through Datadog Workflow Automation rather than a plain Slack notification. Every new or regressed env:prod Error Tracking issue triggers the Error Tracking Triage workflow (handle error-tracking-triage, PPL-3274), which dedupes on the Datadog Issue ID field, files a PPL ticket or updates the open one (a regression on a Done or Closed ticket gets a new linked ticket instead), posts once to #tech-errors at P2/P3, and sets the Datadog issue to ACKNOWLEDGED so that group stops re-alerting. A circuit breaker suppresses auto-creation after five tickets in fifteen minutes. What priority a ticket then gets, and how the pipeline fits together, is on the error severity and triage page.

The monitor reaches the workflow only by naming it: error_tracking_notification_target is @workflow-error-tracking-triage(issue_id="{{issue.id}}"), not a Slack handle. Only a workflow that has a monitor trigger can be named this way, and a handle that resolves to nothing notifies nothing and reports no error, so confirm it in the monitor's own workflow picker after changing it. The handle replaces the Slack one rather than joining it: the workflow sends its own #tech-errors message, so keeping both would post every error twice.

The issue_id argument is not optional. A monitor-triggered run receives an event with no group_value, so the workflow cannot read the alerting issue from the alert itself, and its declared issue_id input has no default. Name the workflow without it and every run fails in under a second with input parameter "issue_id" is required, before the first step. That is silent unless someone looks at the workflow's execution history: the monitor reports a successful notification either way.

Terraform owns the workflow (create_error_tracking_workflow). The original was built by hand during PPL-3274's dry run, so the first apply carrying that flag (release 1.1.34) planned a create and failed against the existing one; the hand-built copy was deleted rather than imported, leaving Terraform to create it from this configuration. Nothing was lost in the swap because the committed spec and the live one matched step for step, but two things do change on any recreation: the workflow ID, which invalidates existing links to it, and its runAsUser, which becomes whoever owns the credentials that created it. That second one matters at run time, because the Jira and Slack steps act as that user.

The Jira JQL calls authenticate differently from the Jira actions, and fail silently. Create_Issue and Add_Regression_Comment are native jira.* actions that go through the Datadog Jira integration account. The two JQL searches (Search_Dedupe, Circuit_Breaker_Search), the two label PUTs and the Link_Regression_Ticket POST cannot: there is no native Jira search action, so they are raw http.request steps carrying Basic Auth from a hand-made Datadog HTTP connection (datadog_jira_http_connection_id, Datadog > Actions > Connections). The Jira actions working therefore says nothing about whether the searches work.

When that credential is rejected, Jira answers POST /rest/api/3/search/jql with 200 and {"isLast": true, "issues": []}, not 401. The only tell is the response header x-seraph-loginreason: AUTHENTICATED_FAILED. Read as a result, an empty issues array means "no ticket exists", so dedupe files a duplicate for every already-ticketed issue, the entire regression branch behind the dedupe hit becomes dead code, and the circuit breaker can never trip no matter how many tickets a burst creates. That happened for two weeks (PPL-3816): #tech-errors carried only NEW posts and not one REGRESSION, and duplicates such as PPL-3815 were filed.

Assert_Dedupe_Search_Authenticated and Assert_Circuit_Breaker_Search_Authenticated now sit between each search and the branch that reads it, and throw on any non-OK seraph verdict, so a rejected credential fails the run instead of quietly inverting the dedupe decision. An absent header is treated as no verdict rather than a failure, so an authenticated search that legitimately returns zero rows still takes the create branch. Recovering from a trip means re-authenticating the connection; nothing in this repo holds that credential.

That connection authenticates as an Atlassian service account, which shapes the URLs. A service account has no site login, so Basic Auth against percihealth.atlassian.net is refused with 401 however good the token is. Its scoped API token is accepted only through the cloud gateway, so all raw calls address the site by cloud id (local.jira_api_base in workflows.tf) rather than by hostname. The token needs read:jira-work for the two searches and write:jira-work for the label PUTs and the issue-link POST, and the account needs Browse Projects and Link Issues on PPL plus visibility of the Datadog Issue ID field (customfield_10798) -- a field it cannot see makes the JQL match nothing while still answering 200, which is the same silent failure in a different costume. Verify a change with a query that must return a row, never with one whose empty result looks like success.

Because a failed run reaches neither Slack nor Jira, [Workflows] Terraform-managed workflow execution failed (monitors.tf) alerts @slack-tech-alerts on status:failure executions of any workflow tagged managed_by:terraform. Those guards are only loud through that monitor, and it also covers the issue_id failure described above.

Managing the workflow from CI needs Actions API access on the Datadog application key. Without that scope the create fails with actions API access is not enabled on this application key, and because the monitor is in the same apply, a failed run can still leave the monitor rewritten while the workflow is untouched: that is how release 1.1.34 reverted the monitor to Slack-only while leaving auto-ticketing dead.

Parallel run (PPL-3554). Monitor 56231493 (New issue to review in production, notifying #tech-alerts) duplicates the Terraform monitor's query exactly and is deliberately still live as a control channel: the same error should show up in #tech-alerts from the old monitor and in #tech-errors from the workflow, the latter carrying a PPL ticket link. Delete 56231493 once that is confirmed. It is hand-made and absent from Terraform, so it has to go from the Datadog UI.

While it is live it also covers the one hole in the workflow path: a workflow that fails before its Slack step notifies nobody, because the monitor no longer posts to Slack itself. That safety net is now held by [Workflows] Terraform-managed workflow execution failed rather than by 56231493, so deleting 56231493 no longer leaves failed runs unheard.

There is no P1 / paging tier. #tech-alerts-p1 does not exist yet. A P1 tier was drafted and parked as PPL-3325 because, checked against real production data, its proposed criteria (crash rate, session-share impact) would not have fired even once in over a year of env:prod Error Tracking history. Do not assume anything in this system pages someone: if an issue needs to interrupt someone immediately, that has to happen by another route until PPL-3325 is unparked.

Environment-to-reporting matrix

Environment Backend (functions) Members / clinicians apps Web Professionals Reported env tag
Developer machine (flutter run, pnpm run dev, emulator suite) No No No —
CI (unit tests, Patrol, preview builds under test) No No No —
Preview (per-PR deploy) Yes Yes Yes preview-pr-<run id>
Staging Yes Yes Yes staging
Staging (semantics build) n/a Yes n/a staging-semantics
Production Yes Yes Yes prod

staging-semantics is the members app built from config/semantics/environment.json, which is why env:staging-semantics events are expected rather than a sign of mis-tagging.

Medplum bots do not initialise a Datadog client at all, so they have nothing to gate.

How each gate works

Every gate is a positive allowlist and fails closed: an environment that cannot be identified as a deployed one reports nothing. Adding an environment to an allowlist means editing code and shipping it — no configuration value can introduce a new reporting environment on its own.

The environment name still comes from configuration (ENV / DD_ENV for the backend, EnvString for the apps, VITE_DD_ENV for Web Professionals), so writing an allowlisted name into a local config is not physically impossible. What matters is that no committed default does it: every checked-in local config carries a non-reporting value, and the backend additionally demands a GCP-injected FUNCTION_TARGET/K_SERVICE that a .env cannot fake, so nothing reports by accident and there is no debug flag to leave switched on.

Backend — apps/perci-platform-backend/functions/src/utils/datadogReporting.ts. apps/perci-platform-backend/functions/src/index.ts only starts the tracer, via initDatadogTracing in apps/perci-platform-backend/functions/src/observability/datadogTracing.ts, when all of these hold: not running under the emulator, NODE_ENV=production, FUNCTION_TARGET or K_SERVICE present (only GCP sets these, so it proves the process is deployed), ENV is one of prod, dev or preview, and both DD_ENV and DD_VERSION are non-blank. ENV=dev is the staging deploy, which reports as DD_ENV=staging. The GCP marker is what makes the gate safe: a local .env is seeded from config/env/functions.staging.env, so a developer machine legitimately holds ENV=dev and NODE_ENV=production.

A deployed run missing DD_ENV or DD_VERSION reports nothing and logs a warning at startup. Untagged telemetry is worse than none — it lands in Error Tracking with no environment or version, so it can neither be filtered out nor attributed to a release.

Members and clinicians apps — packages/perci-platform-frontend-shared/lib/core/observability/. Both apps' main calls the one shared initializeDatadog (datadog_bootstrap.dart), passing only what differs between them: the DatadogApp identity in each app's lib/core/observability/datadog_app.dart, and the BFF they talk to. That call returns without building any configuration unless the build is a release build and EnvString is prod, staging, staging-semantics, or starts with preview-pr- (shouldReportToDatadog in datadog_reporting.dart), so a debug, profile, simulator or flutter test run constructs no client at all. The checked-in assets/environment_values/environment.json carries EnvString: dev, which is deliberately not in the allowlist; deploys overwrite it from config/<env>/environment.json.

The Patrol harness (patrol_test/patrol_setup.dart) never initialises the SDK, so it builds no configuration either.

All Datadog traffic from these apps — RUM, logs, feature flag assignments and evaluations, and on web the browser SDK bundles themselves — is routed through the app's own BFF (/v1/tel and /v1/ff, see datadog_proxy.dart and apps/perci-platform-backend/functions/src/common/datadog/datadogProxy.ts) rather than datadoghq.eu. Ad blockers and managed-device DNS filtering drop the Datadog hosts by default, and a blocked user is invisible in RUM and silently excluded from every flag rollout. Both apps still load the browser SDK from Datadog's CDN in web/index.html; the BFF copy is the fallback for when that load is blocked.

Web Professionals — web-professionals/src/datadog.ts. initializeDatadog returns early unless VITE_DD_ENV is prod, staging, or a preview-pr- value. It no longer falls back to Vite's MODE, which used to tag local pnpm run dev sessions as env:development. The reported tag is the normalised value, so prod and " PROD " cannot show up as two environments.

Reading Error Tracking issues

Two things are easy to misread when triaging:

An issue's title is one sample, not the whole group. Error Tracking groups by stack, not by environment, so a single issue can hold events from staging and production at once and show whichever sample it picked. Issue 56a32128-268f-11f1-8412-da7ad0900005 is titled Fetch timed out ... https://api.staging.perci-fhir.com/oauth2/token yet legitimately matches an env:prod search: over 30 days its staging events hit api.staging.perci-fhir.com while its production events hit contenthub.percihealth.com, api.uk.awellhealth.com and api.cal.com. No production service calls a staging host — that was checked directly and every error mentioning api.staging.perci-fhir.com is env:staging in project perci-platform-staging. Filter or group by env before reading a message as evidence about production.

Backend versions are commit SHAs by design. DD_VERSION is set to github.sha for both staging and production deploys, so a bare SHA in last_seen_version is normal for perci-platform-backend and is not a sign of non-prod telemetry. The Flutter apps report <pubspec version>-<build number> instead.