Skip to content

Error severity and triage

Every production error that reaches Error Tracking becomes a PPL ticket, and every ticket needs a priority. This page is the canonical definition of that priority: what counts as urgent, what counts as ordinary, and what is not an error at all. It is written to be mechanically evaluable, so that two engineers reach the same answer unprompted and a Datadog workflow can compute the same answer without a human in the loop.

It is the reference for the rest of the triage pipeline (epic PPL-3265), and it is also the record of where the pipeline as built differs from the rubric as originally specified. Some of those differences are deliberate: a criterion that could not be justified against real production data. Others are open defects in the live pipeline. Both kinds are listed in Known gaps between this page and the gate.

Three neighbouring pages own the questions this one does not:

  • Error taxonomy: which backend errors are expected 4xx business outcomes rather than defects. Anything classed as an expected outcome there should never reach Error Tracking, so it should never reach this page.
  • Datadog error triage: which environments are allowed to report at all, how the monitor reaches the workflow, how the workflow authenticates to Jira, and how to read an Error Tracking issue without misattributing it.
  • Delivery: the PPL statuses and resolutions this page maps from.

Source of truth: the observed prod corpus

The thresholds on this page were derived from the live prod Error Tracking corpus and RUM session volume, pulled on 5 August 2026. The derivation is shown in How the thresholds were chosen so a future reader can re-derive or challenge them rather than taking them on trust. Re-run it when traffic changes materially. PPL-3272 has since reset the prod corpus to an announced cutoff, so these figures describe the pre-reset history.

What is live today

Checked against production on 2 October 2026.

Piece Ticket State
Issue-level Error Tracking monitor, one notification per new or regressed env:prod issue PPL-3273, PPL-3554 Live, Terraform-managed. Notifies the triage workflow only, not Slack
Inbound triage workflow: dedupe, create or update a PPL ticket, post to #tech-errors, set Datadog state PPL-3274 Live. The Terraform-managed workflow has been published since 25 August 2026 (157 runs to 2 October, 13 failed); a hand-built predecessor filed tickets from 13 August. 156 tickets filed in all. One open defect, see below
Workflow failure monitor, [Workflows] Terraform-managed workflow execution failed PPL-3816 Live, posts to #tech-alerts
Outbound Jira automation: ticket status drives Datadog state PPL-3275 Live, three rules on project PPL
Error count fields on the ticket (Error Occurrences, Impacted Sessions, Error Last Seen) PPL-3352 Live, stamped at creation only. Nothing refreshes them
P1 / paging tier PPL-3325 Parked (Jira status Blocked), deliberately. Nothing pages anyone
Criticality reference table none yet Not built. Unpark condition for PPL-3325
Weekly triage rotation PPL-3276 Not started

The monitor is error-tracking("env:prod").source("all").new().rollup("count").by("issue.id").last("1d") > 0 at monitor priority 3, grouped by issue.id so each issue notifies once. new() covers both a genuinely new issue and one Datadog reopened on a newer version, which is why the alert body reads the reason off {{ issue.alert_reason }} rather than assuming. Its only notification target is @workflow-error-tracking-triage(issue_id="{{issue.id}}"), so every #tech-errors message comes from the workflow, not from the monitor. A second monitor, [Backend] HTTP 4xx elevated, watches 4xx log volume at monitor priority 4 and is not part of this rubric.

A live defect: every ticket arrives Medium

The severity gate never fires. The gate reads impacted_sessions and is_regression from the workflow's Get_Issue_Metrics step, a searchIssues call. In production that step returns total_count and nothing else, so both conditions are always false. Every one of the 131 tickets the workflow has filed since 2 September 2026 arrived Medium (one was raised to Highest by hand afterwards), and Impacted Sessions reads 0 on every ticket that carries the field, including issues that Datadog reports with 10 or more impacted sessions. Until this is fixed, all prioritisation is manual (PPL-4168). The one exception is a regression on a Done or Closed ticket: the workflow knows it is a regression from Jira, so its new ticket is High.

Fixed: regressions failed from 30 September to 2 October 2026

The regression path used to move the existing ticket to Open. On 30 September 2026 PPL lost its Open status (To Do is now Triage, Backlog and Ready), so Jira answered Cannot transition issue to status. and every regression run stopped before its label, comment, Slack post and ACKNOWLEDGED write. The workflow no longer transitions tickets at all, see step 3 below (PPL-4169).

Not an error

Before triaging, check that the thing is a defect at all. These are fixed at source, not prioritised:

Category Example Where it is fixed
Expected 4xx business outcome Wrong password, no care-team allocation, expired session Error taxonomy: logged at warn, kept out of Error Tracking
User input rejection ValidationError on a malformed form field As above
Dependency failure absorbed by a retry A Firestore ABORTED contention that the retry then succeeds on Only recorded once retries are exhausted (PPL-3269)
Telemetry from a non-production environment A preview build tagged env:prod Datadog error triage: environment allowlists
Browser noise nobody can act on Third-party console errors, dropped WebSockets, camera permission denied Proposed as a monitor-query exclusion in PPL-3941, not yet live

If one of these appears in Error Tracking, the bug is in the classification, not in the flow that threw it. Raise a ticket against the taxonomy and set the Datadog issue to EXCLUDED, not IGNORED (see Datadog states).

Three tiers of knowledge

The rubric mixes three kinds of signal, and keeping them apart is what makes the automatable part automatable.

Tier What it is Who evaluates it
Datadog knows natively impacted_sessions, total_count, regression flag, first_seen, first_seen_version, service, error_type The workflow, off the issue object. Today only total_count and the identifying fields actually reach it, see the defect above
We teach Datadog Whether the code path is clinically critical Nobody yet: the criticality reference table does not exist
Not automatable "Is this a repeatable functional break?", "is this cosmetic?" A human, at the weekly triage

Because the third tier cannot be computed, the workflow assigns the highest priority whose mechanical criteria match, and defaults to the lowest when none match. Weekly triage promotes or demotes from there. A ticket that arrives Medium and should have been High is normally a triage correction, not an automation bug, and those corrections are the tuning signal for this page. While the gate is blind, every ticket is that case.

The rubric

Priorities are Jira priorities, not P-numbers

PPL uses the standard five-value scheme (Highest, High, Medium, Low, Lowest), and the workflow sets that field directly. The epic's P1/P2/P3 language maps onto it as follows, and the Jira value is the one that exists in the tool:

Epic tier Jira priority Status
P1, page immediately Highest Not issued by anything. Parked, see below
P2, ticket this sprint High In the gate, but never issued in practice (see the defect above)
P3, backlog Medium The default, and today the only value the gate issues

The severity gate as built

The whole of the workflow's severity computation is two conditions, in the Compute_Context step of infrastructure/terraform/modules/datadog-observability/workflows.tf:

const metricsAttrs = (metricsIssues[0] && metricsIssues[0].attributes) || {};
const impactedSessions = metricsAttrs.impacted_sessions ?? 0;
const isRegression = metricsAttrs.is_regression === true || regressionOfKey !== "";

const SESSIONS_THRESHOLD = 10;
const priorityName = isRegression || impactedSessions >= SESSIONS_THRESHOLD ? "High" : "Medium";
Priority Mechanical test
High The issue carries a regression flag, or its newest PPL ticket is Done or Closed, or impacted sessions >= 10
Medium Everything else

Both inputs come from Get_Issue_Metrics, a searchIssues call for issue.id:<id> windowed from the issue's first_seen to now. PPL-3451 moved the regression read there on the evidence that the getIssue payload has no is_regression attribute. The live step output shows the search returns no such attribute either: a run on 1 October 2026 for an issue with 707 occurrences returned {"total_count": 707} and nothing else. So impactedSessions falls back to 0, the metrics half of isRegression is false, and the gate answers Medium for every ticket except a regression on a finished one, where regressionOfKey comes from Jira instead. Any fix has to request the impact and regression fields explicitly, or read them from a source that returns them, and then be checked against a real run's step output rather than against the corpus pull.

Two further properties of the window, which matter once the inputs arrive:

  • For a newly-seen issue, first_seen to now is close to the 24-hour window the thresholds were derived for. For a regression reopened months later it spans the issue's whole history, so the count is cumulative, but a regression is already High on the first condition.
  • The search has no env filter, and an Error Tracking issue groups across every environment that reports it, so the counts include preview and staging traffic. PPL-3712 (#2160, open) scopes it to env:prod.

Every issue gets a ticket and a #tech-errors message regardless of priority, with two exceptions: the circuit breaker, and a run that fails before its Slack step. Severity sets the priority and the wording of the Slack message; it never decides whether the work is recorded.

What the workflow does per issue

  1. Fetch the issue (getIssue), then fetch its metrics (searchIssues).
  2. Search PPL for an existing ticket by "Datadog Issue ID" ~ <issue id>, then fail the run if Jira did not authenticate the search. Jira answers an unauthenticated JQL search with 200 and no results, which would otherwise read as "no ticket" (PPL-3816, see Datadog error triage).
  3. Open ticket found (newest ticket by created, in a To Do or In Progress status): leave its status alone, add the regression label, comment with the version it returned in, post REGRESSION to #tech-errors, set the Datadog issue to ACKNOWLEDGED.
  4. Finished ticket found (Done or Closed): the PPL workflow has no transition from either back to a To Do status, so the ticket cannot be reopened. Take the create path in step 5, with three differences: the new Bug is High, its description starts with Regression of <old key>, and it gets the regression label and a relates to link to the old ticket. The Slack post reads REGRESSION of <old key>. Because dedupe takes the newest ticket, the next regression of the same issue updates this new ticket. A failed link falls through to the label, and a failed label to the Slack post.
  5. No ticket found: check the circuit breaker, then create a Bug in PPL and post NEW to #tech-errors, then set the issue to ACKNOWLEDGED.

The ticket the create path files carries:

Field Value
Summary [<service>] <error type>: <first line of the message>, cut at the first { and capped at Jira's 254 characters
Description Datadog issue link, service, error type, first-seen and last-seen time and version, total occurrences, impacted sessions, and the full message in a {code} block
Priority High or Medium from the gate
Component From the service to component mapping
Datadog Issue ID, Error Occurrences, Impacted Sessions, Error Last Seen Stamped once, at creation
Reporter The Jira integration account

A run that fails at any step reaches neither Slack nor Jira past that step, and leaves the Datadog issue OPEN. The [Workflows] Terraform-managed workflow execution failed monitor is the only thing that makes that visible.

P1 is parked, and nothing pages anyone

There is no P1 tier and no #tech-alerts-p1 channel. PPL-3325 holds that work, and it is parked on evidence, not on capacity: checked over a 400-day env:prod window, the originally drafted P1 criteria would have fired zero times.

Original P1 criterion Verdict
is_crash == true Dead. No crashes exist. Every issue is platform: BROWSER or BACKEND, and a crash-filtered search returns no data. Meaningful only if native iOS or Android builds start reporting
Impacted sessions as a share of active sessions in 15 minutes Dead. Impact accumulates over weeks rather than spiking; see why the rate expression fails
Blocks a clinical flow Viable, but depends on the criticality reference table, which does not exist
New issue on a release less than 2 hours old Viable and demonstrably matchable, since first_seen_version ties to a real release

When PPL-3325 is unparked, the direction is to rebase on what actually distinguishes urgency in this data: clinical-flow failure at any volume, a new issue on a fresh release, a rate change against the issue's own baseline, and possibly an absolute count over a longer window (around 50 impacted sessions in 24 hours would have caught exactly the top five issues in 14 months and nothing else). Do not reinstate the crash or 15-minute-share criteria without new evidence that they can match.

Nothing in this pipeline interrupts anybody

The error monitors are notifications at monitor priority 3 and 4. Error messages land in #tech-errors and workflow failures in #tech-alerts. If an error genuinely needs someone now, escalate by hand, in the same way as any other unplanned production problem. Do not assume an alert reached a human because a ticket exists.

High: ticket this sprint

Mechanically, either of:

# Signal Test
1 Regression The issue carries a regression flag, meaning Datadog reopened a RESOLVED issue on a newer version
2 Sustained session impact Impacted sessions >= 10

Neither reaches the gate today, so apply both by hand at triage: the Datadog issue page shows the regression and the impacted-session count, which the ticket's Impacted Sessions field does not.

And by human judgement at triage, promote to High when the error is a repeatable functional break: a user can reproduce it by following a documented flow. That is the criterion no workflow can evaluate, and it is the main reason the weekly triage exists.

The derivation below produces a 3 impacted sessions figure as the natural boundary for this tier. The gate does not implement it; it uses a single threshold of 10 for both. Treat 3 as the guide when triaging by hand, and see the gaps table.

Medium: backlog

The default. Everything that matched no High criterion, plus anything triage judges to be cosmetic, single-user, non-blocking, or self-recovering (it degrades then recovers without user action).

Known gaps between this page and the gate

Read this before concluding that the automation is wrong or that this page is stale. Each row is a deliberate divergence or an open defect, not drift.

Rubric criterion Gate as built Why, or what to do
Impacted sessions >= 10 implies High Implemented, never fires searchIssues returns no impacted_sessions, so the value is always 0. Open defect, PPL-4168
Regression implies High Fires only for a regression on a Done or Closed ticket PPL-3451 moved the read to searchIssues, which returns no is_regression either; the finished-ticket case is known from Jira instead. Open defect for the rest, PPL-4168
A regression reopens its ticket Updates an open ticket, files a new linked Bug for a finished one No Done or Closed ticket has a transition back to a To Do status, so a finished ticket is never reopened. PPL-4169
Counts describe prod Counts span every environment The metrics search has no env filter. PPL-3712, #2160 (open)
Expected 4xx outcomes get no ticket Some still do dd-trace marks the middleware span whatever the handler answers. Root cause fixed in PPL-3661; PPL-3712 adds a monitor exclusion as a backstop
is_crash == true implies P1 Not implemented No crash has ever been recorded. Parked with PPL-3325
Clinical-flow failure implies P1 Not implemented Needs the criticality table, which needs Product sign-off
New issue on a release under 2 hours old implies P1 Not implemented Viable, parked with PPL-3325. Needs a deploy-timestamp lookup, which is not on the issue object
Session share over 15 minutes Not implemented Not viable at our traffic, and the monitor uses new() with no rate term at all
High at 3 impacted sessions in 24 h Single threshold of 10, over first_seen to now Worth reconciling once the gate can see sessions at all: either lower the constant to 3, or restate the tier at 10 and drop the 3
Every new issue gets a ticket True except under the circuit breaker or a failed run By design for the breaker, see issues that get no ticket

How the thresholds were chosen

Pulled 5 August 2026 over a trailing 30-day window (env:prod, so roughly 6 July to 5 August 2026), paging the whole corpus rather than sampling the top of the list.

Corpus shape. 150 prod Error Tracking issues, which reconciles exactly:

Service Issues Carry impacted_sessions Events
clinicians-app-frontend 50 50 35,675
members-app-frontend 40 40 6,071
perci-platform-backend 60 0 14,800

The 60 backend issues carry no impacted_sessions attribute at all: they are platform: BACKEND, which has no session concept. Any session-based threshold is therefore frontend-only, and with the crash, criticality and fresh-release criteria all parked, backend severity rests on the regression flag alone until PPL-3325 is picked up. The percentiles below are over the 90 frontend issues that have the attribute, because computing a percentile over 150 values where 60 are absent would be arithmetic on a hole.

Distribution of impacted_sessions across the 90 frontend issues (30 days):

Statistic Value
min / median / max 1 / 2 / 280
mean 16.0
p50 / p75 / p90 2 / 8 / 18
p95 / p99 105 / 280

The distribution is extremely top-heavy: the 5 largest issues hold 67.9 % of all impacted sessions, and the top 3 hold 50.1 %. The upper tail has a clean break, with values running 18, 20, 39, then jumping to 57, 105, 151, 165, 276, 280.

Converting to a per-day rate. impacted_sessions is measured over whatever window you query, so a 30-day figure is cumulative and cannot be compared to a 24-hour threshold directly. Dividing by the 30-day window:

Issue rank 30-day impacted sessions Mean per day
1 280 9.3
2 276 9.2
3 165 5.5
4 151 5.0
5 105 3.5
6 57 1.9

This gives both numbers the epic asked for:

  • 10 impacted sessions in 24 hours as the urgent threshold. No issue in the corpus averages more than 9.3 per day, so nothing then in prod would have tripped it on volume alone. That was the intent: the corpus contained no acute outbreak, only slow burns accumulated over a month. On an app serving roughly 70 to 120 sessions a day, 10 affected sessions in one day is a blast radius of order 10 %. With P1 parked, this is the constant the gate now uses for High.
  • 3 impacted sessions in 24 hours as the boundary for ordinary-but-real impact. Exceeded by 5 of 90 issues on a per-day basis, so about the 94th percentile of observed per-day rates: high enough that the median issue (0.07 sessions per day) never reaches it, low enough to catch a real functional break early. The gate does not use it; triage should.

Why 24 hours was the natural window, and what the workflow actually does

The Datadog issue-detail endpoint takes no time parameters: its figures always cover the 24 hours before the issue was last seen, so a 24-hour count comes free with fetching the issue. The shipped workflow instead runs a searchIssues call from the issue's first_seen to now, which for a new issue is close to the same number and for an old one is cumulative. Both thresholds above are stated against 24 hours because that is the window they were derived for. As called today, that search returns no impacted_sessions at all, so neither threshold is evaluated (see the severity gate as built).

The rate expression is not viable at current traffic

The epic asked for two expressions of the session criterion: a rate against active sessions for the monitor, and an absolute count for the workflow. The absolute count is above. The rate, as originally specified over a 15-minute window, does not work at our traffic level, and the number matters more than the shape of the rule.

Prod RUM sessions (@type:session, @session.type:user) on 21 July 2026, the busiest day in the retained window, in 15-minute buckets:

App Sessions that day Non-empty 15-min buckets Peak bucket Median non-empty bucket
Member App 122 53 of 96 6 2
Clinical App 71 44 of 96 4 1

At a peak of 6 concurrent sessions per 15 minutes, one session is 16.7 % of the busiest window on the busiest day, and 50 % of a median window. Any percentage threshold is therefore decided by a single user: 10 % of the peak bucket is 0.6 sessions, so the threshold rounds down to "one error". A rule like that looks quantitative and discriminates nothing.

This is why PPL-3273 shipped a monitor that counts new issues rather than a share of sessions, and why the share-based P1 criterion is parked rather than re-tuned. The best rate expression available at this traffic level, if one is ever wanted, is a 6-hour window at 20 % of active sessions with a floor of 5 impacted sessions: at roughly 5 sessions an hour on a weekday, 6 hours accumulates 25 to 30 sessions per app, the smallest denominator for which 20 % still requires 5 distinct impacted sessions and cannot be tripped by one user. The floor is what keeps the quiet hours safe.

When to revisit the 15-minute window. For a 20 % threshold to require 5 impacted sessions within 15 minutes, a single app needs about 25 active sessions per 15 minutes, which is roughly 100 an hour or 2,400 a day. Current peak is 6 per 15 minutes. So a 15-minute rate becomes meaningful at roughly 20 times today's traffic. Re-run the bucket query above before adopting it.

Clinical criticality (proposal)

Not built, and not signed off

The flow list below is a proposal from engineering, seeded from the flows named in PPL-3271 plus the view paths and backend modules that actually carry prod errors. Product has not agreed it, no reference table exists, and PPL-3274 shipped without one. It is now an explicit unpark condition for PPL-3325: clinical-flow detection is the criterion that matters most for a health platform and the one with the least data behind it.

Service-level tiering is useless here: there are three services and all three matter. The signal has to be at flow level, so criticality belongs in a Datadog Reference Table, keyed on the service plus the code location, so Product can change what counts as critical by editing data rather than by editing a workflow.

Intended table shape:

Column Type Meaning
service string members-app-frontend, clinicians-app-frontend, perci-platform-backend
match_type string view_path, file_path, or function_name
match_value string The value to match, exactly
criticality string clinical-critical or standard
flow string Human-readable flow name, used in the ticket summary and the Slack message

Primary key is (service, match_type, match_value). A lookup that misses returns standard, so the table is a positive allowlist and an unmapped path never silently becomes urgent.

Proposed clinical-critical rows. The view paths are the real paths carrying prod errors today; the backend files are the real modules that produced prod issues.

Flow Service match_type match_value
Joining a call clinicians-app-frontend view_path /appointmentCall
Joining a call members-app-frontend view_path /videoAppointment
Joining a call perci-platform-backend file_path src/bff-clinical/appointments/endAppointmentCall.ts
Appointment conclusion clinicians-app-frontend view_path /appointmentConclusion
Appointment conclusion clinicians-app-frontend view_path /appointmentConclusionSummary
Appointment conclusion perci-platform-backend file_path src/bff-clinical/appointments/upsertAppointmentConclusion.ts
Referral intake members-app-frontend view_path /referralSignup
Referral intake perci-platform-backend file_path src/bff/referrals/verifyReferralTokenData.ts
Login members-app-frontend view_path /login
Login members-app-frontend view_path /start-sso
Login members-app-frontend view_path /sso-callback
Login clinicians-app-frontend view_path /signin
Login perci-platform-backend file_path src/bff/middleware/descopeMemberMiddleware.ts
Login perci-platform-backend file_path src/bff-clinical/middleware/descopeClinicianMiddleware.ts

Candidates engineering suggests Product also considers, because they carry prod errors and plausibly block care rather than merely inconveniencing a user: appointment booking (/bookAnAppointment, src/bff/appointments/createAppointment.ts), the code and SSO signup chain (/ssoSignupHome, /codeSignupAbout, /codeSignupBasicInfo), and clinical messaging (/MessagesHome, /newMessagesHome). These are not in the proposed list above; they are the agenda for the Product conversation.

Backend file_path values need normalising

Backend issues report file_path with the container prefix, for example /workspace/src/bff/referrals/verifyReferralTokenData.ts. Strip the /workspace/ prefix before the lookup, or store the prefixed form. Note also that 16 of the 60 backend issues report a file_path inside node_modules (13 of them in @grpc/grpc-js), which no criticality row will ever match. Those fall through to standard correctly.

Datadog states

Five states exist in the public API. These are the only values PUT /api/v2/error-tracking/issues/{issue_id}/state accepts:

State Meaning here
OPEN Never triaged. Every issue the workflow handles successfully is moved to ACKNOWLEDGED, so an OPEN issue that has alerted means the workflow run failed. Today that is mostly regressions, see the live defects
ACKNOWLEDGED Recorded and being worked, or suppressed by the circuit breaker
RESOLVED Fixed and released to production. Reopens automatically if it recurs on a newer version
IGNORED Deliberately never actioned. Stays silent forever, including on new versions
EXCLUDED Not a real error, should never have been recorded. Use this for taxonomy and environment-tagging mistakes, not for "we chose not to fix it"

FOR_REVIEW and REVIEWED are not API values

Some tooling, including the Datadog MCP server and parts of the UI, uses FOR_REVIEW and REVIEWED as aliases for OPEN and ACKNOWLEDGED. They are aliases only. Sending either to the state endpoint fails. When reading state from a tool that returns them, map FOR_REVIEW to OPEN and REVIEWED to ACKNOWLEDGED before comparing against anything on this page.

Why RESOLVED and IGNORED are not interchangeable

This distinction is the entire basis of regression detection, and it is the one state decision worth getting right:

  • A RESOLVED issue that recurs on a newer version reopens itself and is flagged as a regression. That is free regression detection, with no memory and no spreadsheet.
  • An IGNORED issue never comes back, on any version, ever.

So IGNORED is not "closed"; it is "permanently deaf to this error". Reach for it only when the error is genuinely acceptable forever. If there is any chance the fix matters later, RESOLVED is the safer choice: the worst case is one reopened ticket, whereas a wrongly ignored issue is a silent defect with no alarm attached.

Datadog also resolves issues on its own, too early

Datadog's GitHub integration sets an issue RESOLVED the moment a PR it has linked to the issue merges into develop, days before the fix reaches production, and whether or not the PR fixed anything. An occurrence in that gap reopens the issue as a false regression. The fix is a Datadog console change (turn off PR-based auto-resolve), tracked on PPL-3711. Until then, a RESOLVED issue does not prove the fix is released.

For scale, the 5 August corpus was 98 OPEN, 12 ACKNOWLEDGED, 37 RESOLVED and 3 IGNORED, and 59 of the 150 issues already carried a regression flag. State then was close to meaningless, which is the problem the pipeline exists to fix.

Issues that get no ticket

The workflow carries a circuit breaker, and it is the one designed case where an issue is ACKNOWLEDGED with no ticket behind it. Before creating a ticket it counts PPL tickets with a non-empty Datadog Issue ID created in the last 15 minutes; if there are more than 5, it posts a Circuit breaker tripped message to #tech-errors, sets the issue to ACKNOWLEDGED, and creates nothing.

That protects the board from an error storm at the cost of losing the tickets for it. So a circuit-breaker message in #tech-errors is itself a triage action: work out what broke, raise one ticket for it by hand, and check Error Tracking for ACKNOWLEDGED issues from that window with no Datadog Issue ID recorded against them.

The other case is undesigned: a run that fails. It leaves the issue OPEN with no ticket, or with an untouched ticket on the regression path, and announces itself only through [Workflows] Terraform-managed workflow execution failed in #tech-alerts.

Ticket status to Datadog state

Three Jira automation rules on project PPL (PPL-3275), all gated on Datadog Issue ID being non-empty, each a Send web request to PUT https://api.datadoghq.eu/api/v2/error-tracking/issues/{issue_id}/state with a service-account app key scoped to error_tracking_read and error_tracking_write only. Done means released to production in this project (see Delivery), which is why Done maps to RESOLVED rather than to something earlier. The release chain gets there on its own: a release PR merging to main moves its tickets from Ready for Release to Done, and the Done rule fires.

Rule (as named in Jira) Trigger Datadog state
Set by the workflow, not a rule Ticket created or updated by the workflow ACKNOWLEDGED
Datadog: Done -> RESOLVED Transition to Done RESOLVED
Datadog: Closed as Won't Do or Declined -> IGNORED Transition to Closed, and resolution in ("Won't Do", "Declined") IGNORED
(no rule) Transition to Closed with any other resolution Left as it was, normally ACKNOWLEDGED
Datadog: reopened out of Done -> ACKNOWLEDGED Transition out of Done to any status ACKNOWLEDGED

Each rule has "delay execution until web request response" enabled, so a failure shows up in the automation audit log instead of dropping silently. Rule-editing on PPL is restricted, because Jira Automation has no secret store: anything in a web-request header is readable by anyone who can edit the rule.

Which resolutions count as won't-do

Project PPL had exactly six resolution values when this was checked. Counts are live usage across the project as of 5 August 2026, and confirm the list was complete (529 + 55 + 47 + 24 + 22 + 1 = 678, which equalled project = PPL AND resolution IS NOT EMPTY).

Resolution Issues Should map to IGNORED? Live rule
Done 529 No. Fixed, so RESOLVED Not IGNORED
Duplicate 55 No. See below Not IGNORED
Won't Do 47 Yes IGNORED
Cannot Reproduce 24 No. Unexplained, not accepted Not IGNORED
Functioned As Designed 22 Yes Not IGNORED. The rule omits it
Declined 1 Yes IGNORED

Won't Do, Functioned As Designed and Declined are the won't-do style decisions. The live rule covers only Won't Do and Declined, so an error closed as Functioned As Designed stays ACKNOWLEDGED and alerts again only as a regression. Add it to the rule's JQL condition, or change this table if leaving it out is intended.

A duplicate must never set IGNORED

Closing a ticket as Duplicate says nothing about whether the error is being fixed, only that this ticket is not where the fix lives. Setting IGNORED would silence an error that another open ticket is still working on, and it would silence it permanently. Leave a duplicate at ACKNOWLEDGED and let the surviving ticket drive the state. Same reasoning for Cannot Reproduce: an error we could not reproduce is unexplained, not accepted, and it should be allowed to come back and tell us more.

The resolution may arrive after the transition

In PPL the Closed and No testing needed transitions carried no resolution screen when checked: getTransitionsForJiraIssue returned empty fields for every transition. Resolutions were set as a separate field edit afterwards. In PPL-2929, for instance, the status moved to Done and the resolution was set two seconds later as its own changelog entry. The IGNORED rule is triggered by the transition and evaluates its resolution JQL at that moment, so if the resolution lands afterwards the condition fails and the rule does nothing. Set the resolution in the same action as the close, or check the automation audit log after closing. No auto-filed ticket has yet been closed as Won't Do or Declined, so the rule has not been exercised in production. Adding resolution screens to those transitions would be the clean fix.

Service to component mapping

Tickets the workflow creates carry the Component matching the Datadog service, so they route on the PPL board like any other ticket (see Delivery).

Datadog service PPL Component Component id
clinicians-app-frontend Clinical App 10075
members-app-frontend Member App 10074
perci-platform-backend Backend 10330

All three are mapped explicitly. The workflow first shipped with only clinicians-app-frontend in the map, so member app errors filed under Backend until PPL-3450 added the Member App row. An unmapped service still defaults to Backend on purpose, so a service nobody has mapped yet lands somewhere a human sees it. web-professionals is one such service today: it reports to prod Error Tracking and files under Backend (for example PPL-4035).

The Datadog Issue ID field

The join key between a Datadog issue and its PPL ticket, added by PPL-3270.

Property Value
Field id customfield_10798
Name Datadog Issue ID
Type Text Field (single line)
Scope Project PPL, optional, on the Bug, Story and Task screens
JQL "Datadog Issue ID" ~ "<uuid>"
Saved filter project = PPL AND "Datadog Issue ID" IS NOT EMPTY

The workflow stamps it on creation and uses it three ways: to find the existing ticket when an issue reopens rather than creating a second one, to count recent auto-created tickets for the circuit breaker, and to scope the PPL-3275 automation rules so they never fire on a hand-written ticket.

On 2 October 2026, 172 PPL tickets carried it: 156 filed by the workflow and 16 stamped by hand. Ten Datadog issue ids appear on more than one ticket. Seven are pairs of hand-stamped tickets from before the workflow. The other three are workflow duplicates filed while the dedupe search was unauthenticated (PPL-3816, fixed 15 September): one issue five times in August, and two issues twice each in early September. Dedupe asks Jira for the newest match (ORDER BY created DESC), so for those ids the regression path acts on the latest ticket.

Frequency and impact on the ticket

PPL-3352 added three custom fields that the workflow stamps when it creates a ticket:

Field Id Source
Error Occurrences customfield_10909 total_count from the metrics search
Impacted Sessions customfield_10910 impacted_sessions from the metrics search, so always 0 today (see the defect)
Error Last Seen customfield_10911 The issue's last_seen

They are a creation-time snapshot. Nothing refreshes them: the triage workflow is the only Datadog workflow, and it writes these fields only on the create path. An error that arrives as 12 occurrences and grows to 4,000 still reads 12 on its ticket months later, which is the difference between a backlog you can prioritise and a pile of equally-sized cards.

The intended end state is a scheduled refresh over open error-derived tickets, so the triage queue becomes a saved filter sorted by Impacted Sessions descending, and a ticket triaged Medium months ago that has quietly become the highest-impact issue on the board becomes visible without anyone remembering to re-check it. That needs both the refresh job and a metrics search that returns impacted_sessions. Until then, read impact off the Datadog issue page, not the ticket.

The weekly triage

A half-hour slot, one owner per week, rotating round-robin through the engineers on the @percihealth/backend and @percihealth/frontend teams (the same two teams that drive PR review routing). PPL-3276 starts the rotation and retires the manual path. It has not started.

The owner works one saved filter, not the Datadog UI. New auto-filed tickets land in Triage:

  1. Set or correct the priority. Every ticket the workflow created since the last session. While the gate is blind every one arrives Medium, so this step is where all High decisions are made: open the Datadog issue link, apply the High criteria by hand, and add the judgement criteria (repeatable functional break, cosmetic) that the gate can never evaluate. Until the criticality table exists, it is also where a clinical-flow failure gets the priority no automation gave it.
  2. Clear anything untriaged. Nothing may stay untriaged for more than seven days. If a ticket has had no priority confirmed within seven days of creation, it is the triage owner's job to give it one that week, or to close it with a resolution.
  3. Pick up what the pipeline dropped. Any Circuit breaker tripped message in #tech-errors, and any [Workflows] Terraform-managed workflow execution failed alert in #tech-alerts, since the last session. Both mean issues exist with no ticket, or with a regression nobody was told about. See issues that get no ticket.
  4. Feed the corrections back. A priority that was wrong for a systematic reason is a change to this page or to the gate, not a one-off edit. Raise it.

CONFIRM: rota and slot

The named rota and the day still need confirming with the team, alongside the wider on-call gaps flagged in Team ownership. There is no out-of-hours escalation path to define yet, because nothing pages: see P1 is parked.

Engineers do not open Datadog

Deliberately. The whole pipeline exists so that Error Tracking state is derived, not maintained by hand:

  • Errors arrive as PPL tickets, and are worked in Jira and GitHub like any other ticket.
  • Datadog state follows the ticket status automatically, per the mapping above.
  • Nobody is expected to log in to mark something reviewed or fixed. That expectation is precisely why Error Tracking state was meaningless before the pipeline.

Open Datadog to investigate an error (stack trace, affected versions, session replay, and today its impacted-session count), never to record a decision. If you find yourself changing state by hand, either the ticket is missing or an automation is broken, and both are worth a ticket of their own.

  • Error taxonomy: expected client outcomes versus defects.
  • Datadog error triage: environment reporting gates, how the monitor reaches the workflow, the workflow's Jira authentication, and how to read an Error Tracking issue.
  • Datadog backend observability: how backend telemetry is produced.
  • Delivery: the Jira statuses and resolutions this page maps from.