Datadog error triage¶
Error Tracking is only useful if env:prod means "real users in production". This page records
which environments are allowed to report, and where each gate lives.
Alerting tiers¶
There is currently one tier, and it runs through Datadog Workflow Automation rather than a plain
Slack notification. Every new or regressed env:prod Error Tracking issue triggers the
Error Tracking Triage workflow (handle error-tracking-triage, PPL-3274), which dedupes on the
Datadog Issue ID field, files a PPL ticket or updates the open one (a regression on a Done or
Closed ticket gets a new linked ticket instead), posts once to
#tech-errors at P2/P3, and sets the Datadog issue to ACKNOWLEDGED so that group stops
re-alerting. A circuit breaker suppresses auto-creation after five tickets in fifteen minutes.
What priority a ticket then gets, and how the pipeline fits together, is on the
error severity and triage page.
The monitor reaches the workflow only by naming it: error_tracking_notification_target is
@workflow-error-tracking-triage(issue_id="{{issue.id}}"), not a Slack handle. Only a workflow that
has a monitor trigger can be named this way, and a handle that resolves to nothing notifies nothing
and reports no error, so confirm it in the monitor's own workflow picker after changing it. The
handle replaces the Slack one rather than joining it: the workflow sends its own #tech-errors
message, so keeping both would post every error twice.
The issue_id argument is not optional. A monitor-triggered run receives an event with no
group_value, so the workflow cannot read the alerting issue from the alert itself, and its
declared issue_id input has no default. Name the workflow without it and every run fails in under
a second with input parameter "issue_id" is required, before the first step. That is silent unless
someone looks at the workflow's execution history: the monitor reports a successful notification
either way.
Terraform owns the workflow (create_error_tracking_workflow). The original was built by hand
during PPL-3274's dry run, so the first apply carrying that flag (release 1.1.34) planned a create
and failed against the existing one; the hand-built copy was deleted rather than imported, leaving
Terraform to create it from this configuration. Nothing was lost in the swap because the committed
spec and the live one matched step for step, but two things do change on any recreation: the
workflow ID, which invalidates existing links to it, and its runAsUser, which becomes whoever
owns the credentials that created it. That second one matters at run time, because the Jira and
Slack steps act as that user.
The Jira JQL calls authenticate differently from the Jira actions, and fail silently.
Create_Issue and Add_Regression_Comment are native jira.* actions that go
through the Datadog Jira integration account. The two JQL searches (Search_Dedupe,
Circuit_Breaker_Search), the two label PUTs and the Link_Regression_Ticket POST cannot: there is no native Jira search
action, so they are raw http.request steps carrying Basic Auth from a hand-made Datadog HTTP
connection (datadog_jira_http_connection_id, Datadog > Actions > Connections). The Jira actions
working therefore says nothing about whether the searches work.
When that credential is rejected, Jira answers POST /rest/api/3/search/jql with 200 and
{"isLast": true, "issues": []}, not 401. The only tell is the response header
x-seraph-loginreason: AUTHENTICATED_FAILED. Read as a result, an empty issues array means "no
ticket exists", so dedupe files a duplicate for every already-ticketed issue, the entire regression
branch behind the dedupe hit becomes dead code, and the circuit breaker can never trip no matter how
many tickets a burst creates. That happened for two weeks (PPL-3816): #tech-errors carried only
NEW posts and not one REGRESSION, and duplicates such as PPL-3815 were filed.
Assert_Dedupe_Search_Authenticated and Assert_Circuit_Breaker_Search_Authenticated now sit
between each search and the branch that reads it, and throw on any non-OK seraph verdict, so a
rejected credential fails the run instead of quietly inverting the dedupe decision. An absent header
is treated as no verdict rather than a failure, so an authenticated search that legitimately returns
zero rows still takes the create branch. Recovering from a trip means re-authenticating the
connection; nothing in this repo holds that credential.
That connection authenticates as an Atlassian service account, which shapes the URLs. A service
account has no site login, so Basic Auth against percihealth.atlassian.net is refused with 401
however good the token is. Its scoped API token is accepted only through the cloud gateway, so all
raw calls address the site by cloud id (local.jira_api_base in workflows.tf) rather than
by hostname. The token needs read:jira-work for the two searches and write:jira-work for the
label PUTs and the issue-link POST, and the account needs Browse Projects and Link Issues on PPL plus visibility of the Datadog Issue ID
field (customfield_10798) -- a field it cannot see makes the JQL match nothing while still
answering 200, which is the same silent failure in a different costume. Verify a change with a
query that must return a row, never with one whose empty result looks like success.
Because a failed run reaches neither Slack nor Jira, [Workflows] Terraform-managed workflow
execution failed (monitors.tf) alerts @slack-tech-alerts on status:failure executions of any
workflow tagged managed_by:terraform. Those guards are only loud through that monitor, and it also
covers the issue_id failure described above.
Managing the workflow from CI needs Actions API access on the Datadog application key. Without that
scope the create fails with actions API access is not enabled on this application key, and
because the monitor is in the same apply, a failed run can still leave the monitor rewritten while
the workflow is untouched: that is how release 1.1.34 reverted the monitor to Slack-only while
leaving auto-ticketing dead.
Parallel run (PPL-3554). Monitor 56231493 (New issue to review in production, notifying
#tech-alerts) duplicates the Terraform monitor's query exactly and is deliberately still live as a
control channel: the same error should show up in #tech-alerts from the old monitor and in
#tech-errors from the workflow, the latter carrying a PPL ticket link. Delete 56231493 once that
is confirmed. It is hand-made and absent from Terraform, so it has to go from the Datadog UI.
While it is live it also covers the one hole in the workflow path: a workflow that fails before its
Slack step notifies nobody, because the monitor no longer posts to Slack itself. That safety net is
now held by [Workflows] Terraform-managed workflow execution failed rather than by 56231493, so
deleting 56231493 no longer leaves failed runs unheard.
There is no P1 / paging tier. #tech-alerts-p1 does not exist yet. A P1 tier was drafted and
parked as PPL-3325 because, checked against real production data, its proposed criteria (crash rate,
session-share impact) would not have fired even once in over a year of env:prod Error Tracking
history. Do not assume anything in this system pages someone: if an issue needs to interrupt
someone immediately, that has to happen by another route until PPL-3325 is unparked.
Environment-to-reporting matrix¶
| Environment | Backend (functions) |
Members / clinicians apps | Web Professionals | Reported env tag |
|---|---|---|---|---|
Developer machine (flutter run, pnpm run dev, emulator suite) |
No | No | No | — |
| CI (unit tests, Patrol, preview builds under test) | No | No | No | — |
| Preview (per-PR deploy) | Yes | Yes | Yes | preview-pr-<run id> |
| Staging | Yes | Yes | Yes | staging |
| Staging (semantics build) | n/a | Yes | n/a | staging-semantics |
| Production | Yes | Yes | Yes | prod |
staging-semantics is the members app built from config/semantics/environment.json, which is why
env:staging-semantics events are expected rather than a sign of mis-tagging.
Medplum bots do not initialise a Datadog client at all, so they have nothing to gate.
How each gate works¶
Every gate is a positive allowlist and fails closed: an environment that cannot be identified as a deployed one reports nothing. Adding an environment to an allowlist means editing code and shipping it — no configuration value can introduce a new reporting environment on its own.
The environment name still comes from configuration (ENV / DD_ENV for the backend, EnvString
for the apps, VITE_DD_ENV for Web Professionals), so writing an allowlisted name into a local
config is not physically impossible. What matters is that no committed default does it: every
checked-in local config carries a non-reporting value, and the backend additionally demands a
GCP-injected FUNCTION_TARGET/K_SERVICE that a .env cannot fake, so nothing reports by
accident and there is no debug flag to leave switched on.
Backend — apps/perci-platform-backend/functions/src/utils/datadogReporting.ts. apps/perci-platform-backend/functions/src/index.ts only starts the
tracer, via initDatadogTracing in apps/perci-platform-backend/functions/src/observability/datadogTracing.ts, when all of
these hold: not running under the emulator, NODE_ENV=production,
FUNCTION_TARGET or K_SERVICE present (only GCP sets these, so it proves the process is
deployed), ENV is one of prod, dev or preview, and both DD_ENV and DD_VERSION are
non-blank. ENV=dev is the staging deploy, which reports as DD_ENV=staging. The GCP marker is
what makes the gate safe: a local .env is seeded from config/env/functions.staging.env, so a
developer machine legitimately holds ENV=dev and NODE_ENV=production.
A deployed run missing DD_ENV or DD_VERSION reports nothing and logs a warning at startup.
Untagged telemetry is worse than none — it lands in Error Tracking with no environment or version,
so it can neither be filtered out nor attributed to a release.
Members and clinicians apps — packages/perci-platform-frontend-shared/lib/core/observability/.
Both apps' main calls the one shared initializeDatadog (datadog_bootstrap.dart), passing only
what differs between them: the DatadogApp identity in each app's
lib/core/observability/datadog_app.dart, and the BFF they talk to. That call returns without
building any configuration unless the build is a release build and EnvString is prod,
staging, staging-semantics, or starts with preview-pr- (shouldReportToDatadog in
datadog_reporting.dart), so a debug, profile, simulator or flutter test run constructs no
client at all. The checked-in assets/environment_values/environment.json carries EnvString: dev,
which is deliberately not in the allowlist; deploys overwrite it from
config/<env>/environment.json.
The Patrol harness (patrol_test/patrol_setup.dart) never initialises the SDK, so it builds no
configuration either.
All Datadog traffic from these apps — RUM, logs, feature flag assignments and evaluations, and on
web the browser SDK bundles themselves — is routed through the app's own BFF (/v1/tel and
/v1/ff, see datadog_proxy.dart and apps/perci-platform-backend/functions/src/common/datadog/datadogProxy.ts) rather than
datadoghq.eu. Ad blockers and managed-device DNS filtering drop the Datadog hosts by default, and
a blocked user is invisible in RUM and silently excluded from every flag rollout. Both apps still
load the browser SDK from Datadog's CDN in web/index.html; the BFF copy is the fallback for when
that load is blocked.
Web Professionals — web-professionals/src/datadog.ts. initializeDatadog returns early
unless VITE_DD_ENV is prod, staging, or a preview-pr- value. It no longer falls back to
Vite's MODE, which used to tag local pnpm run dev sessions as env:development. The reported
tag is the normalised value, so prod and " PROD " cannot show up as two environments.
Reading Error Tracking issues¶
Two things are easy to misread when triaging:
An issue's title is one sample, not the whole group. Error Tracking groups by stack, not by
environment, so a single issue can hold events from staging and production at once and show
whichever sample it picked. Issue 56a32128-268f-11f1-8412-da7ad0900005 is titled
Fetch timed out ... https://api.staging.perci-fhir.com/oauth2/token yet legitimately matches an
env:prod search: over 30 days its staging events hit api.staging.perci-fhir.com while its
production events hit contenthub.percihealth.com, api.uk.awellhealth.com and api.cal.com.
No production service calls a staging host — that was checked directly and every error mentioning
api.staging.perci-fhir.com is env:staging in project perci-platform-staging. Filter or group
by env before reading a message as evidence about production.
Backend versions are commit SHAs by design. DD_VERSION is set to github.sha for both
staging and production deploys, so a bare SHA in last_seen_version is normal for
perci-platform-backend and is not a sign of non-prod telemetry. The Flutter apps report
<pubspec version>-<build number> instead.