Error severity and triage¶
Every production error that reaches Error Tracking becomes a PPL ticket, and every ticket needs a priority. This page is the canonical definition of that priority: what counts as urgent, what counts as ordinary, and what is not an error at all. It is written to be mechanically evaluable, so that two engineers reach the same answer unprompted and a Datadog workflow can compute the same answer without a human in the loop.
It is the reference for the rest of the triage pipeline (epic PPL-3265), and it is also the record of where the pipeline as built differs from the rubric as originally specified. Some of those differences are deliberate: a criterion that could not be justified against real production data. Others are open defects in the live pipeline. Both kinds are listed in Known gaps between this page and the gate.
Three neighbouring pages own the questions this one does not:
- Error taxonomy: which backend errors are expected 4xx business outcomes rather than defects. Anything classed as an expected outcome there should never reach Error Tracking, so it should never reach this page.
- Datadog error triage: which environments are allowed to report at all, how the monitor reaches the workflow, how the workflow authenticates to Jira, and how to read an Error Tracking issue without misattributing it.
- Delivery: the PPL statuses and resolutions this page maps from.
Source of truth: the observed prod corpus
The thresholds on this page were derived from the live prod Error Tracking corpus and RUM session volume, pulled on 5 August 2026. The derivation is shown in How the thresholds were chosen so a future reader can re-derive or challenge them rather than taking them on trust. Re-run it when traffic changes materially. PPL-3272 has since reset the prod corpus to an announced cutoff, so these figures describe the pre-reset history.
What is live today¶
Checked against production on 2 October 2026.
| Piece | Ticket | State |
|---|---|---|
Issue-level Error Tracking monitor, one notification per new or regressed env:prod issue |
PPL-3273, PPL-3554 | Live, Terraform-managed. Notifies the triage workflow only, not Slack |
Inbound triage workflow: dedupe, create or update a PPL ticket, post to #tech-errors, set Datadog state |
PPL-3274 | Live. The Terraform-managed workflow has been published since 25 August 2026 (157 runs to 2 October, 13 failed); a hand-built predecessor filed tickets from 13 August. 156 tickets filed in all. One open defect, see below |
Workflow failure monitor, [Workflows] Terraform-managed workflow execution failed |
PPL-3816 | Live, posts to #tech-alerts |
| Outbound Jira automation: ticket status drives Datadog state | PPL-3275 | Live, three rules on project PPL |
Error count fields on the ticket (Error Occurrences, Impacted Sessions, Error Last Seen) |
PPL-3352 | Live, stamped at creation only. Nothing refreshes them |
| P1 / paging tier | PPL-3325 | Parked (Jira status Blocked), deliberately. Nothing pages anyone |
| Criticality reference table | none yet | Not built. Unpark condition for PPL-3325 |
| Weekly triage rotation | PPL-3276 | Not started |
The monitor is
error-tracking("env:prod").source("all").new().rollup("count").by("issue.id").last("1d") > 0
at monitor priority 3, grouped by issue.id so each issue notifies once. new() covers both
a genuinely new issue and one Datadog reopened on a newer version, which is why the alert body
reads the reason off {{ issue.alert_reason }} rather than assuming. Its only notification
target is @workflow-error-tracking-triage(issue_id="{{issue.id}}"), so every #tech-errors
message comes from the workflow, not from the monitor. A second monitor,
[Backend] HTTP 4xx elevated, watches 4xx log volume at monitor priority 4 and is not part of
this rubric.
A live defect: every ticket arrives Medium
The severity gate never fires. The gate reads impacted_sessions and is_regression
from the workflow's Get_Issue_Metrics step, a searchIssues call. In production that
step returns total_count and nothing else, so both conditions are always false. Every
one of the 131 tickets the workflow has filed since 2 September 2026 arrived Medium (one was
raised to Highest by hand afterwards), and
Impacted Sessions reads 0 on every ticket that carries the field, including issues that
Datadog reports with 10 or more impacted sessions. Until this is fixed, all
prioritisation is manual (PPL-4168). The one exception is a regression on a Done or
Closed ticket: the workflow knows it is a regression from Jira, so its new ticket is High.
Fixed: regressions failed from 30 September to 2 October 2026
The regression path used to move the existing ticket to Open. On 30 September 2026 PPL
lost its Open status (To Do is now Triage, Backlog and Ready), so Jira answered
Cannot transition issue to status. and every regression run stopped before its label,
comment, Slack post and ACKNOWLEDGED write. The workflow no longer transitions tickets at
all, see step 3 below (PPL-4169).
Not an error¶
Before triaging, check that the thing is a defect at all. These are fixed at source, not prioritised:
| Category | Example | Where it is fixed |
|---|---|---|
| Expected 4xx business outcome | Wrong password, no care-team allocation, expired session | Error taxonomy: logged at warn, kept out of Error Tracking |
| User input rejection | ValidationError on a malformed form field |
As above |
| Dependency failure absorbed by a retry | A Firestore ABORTED contention that the retry then succeeds on |
Only recorded once retries are exhausted (PPL-3269) |
| Telemetry from a non-production environment | A preview build tagged env:prod |
Datadog error triage: environment allowlists |
| Browser noise nobody can act on | Third-party console errors, dropped WebSockets, camera permission denied |
Proposed as a monitor-query exclusion in PPL-3941, not yet live |
If one of these appears in Error Tracking, the bug is in the classification, not in the
flow that threw it. Raise a ticket against the taxonomy and set the Datadog issue to
EXCLUDED, not IGNORED (see Datadog states).
Three tiers of knowledge¶
The rubric mixes three kinds of signal, and keeping them apart is what makes the automatable part automatable.
| Tier | What it is | Who evaluates it |
|---|---|---|
| Datadog knows natively | impacted_sessions, total_count, regression flag, first_seen, first_seen_version, service, error_type |
The workflow, off the issue object. Today only total_count and the identifying fields actually reach it, see the defect above |
| We teach Datadog | Whether the code path is clinically critical | Nobody yet: the criticality reference table does not exist |
| Not automatable | "Is this a repeatable functional break?", "is this cosmetic?" | A human, at the weekly triage |
Because the third tier cannot be computed, the workflow assigns the highest priority whose
mechanical criteria match, and defaults to the lowest when none match. Weekly triage promotes
or demotes from there. A ticket that arrives Medium and should have been High is normally a
triage correction, not an automation bug, and those corrections are the tuning signal for this
page. While the gate is blind, every ticket is that case.
The rubric¶
Priorities are Jira priorities, not P-numbers¶
PPL uses the standard five-value scheme (Highest, High, Medium, Low, Lowest), and
the workflow sets that field directly. The epic's P1/P2/P3 language maps onto it as follows,
and the Jira value is the one that exists in the tool:
| Epic tier | Jira priority | Status |
|---|---|---|
| P1, page immediately | Highest |
Not issued by anything. Parked, see below |
| P2, ticket this sprint | High |
In the gate, but never issued in practice (see the defect above) |
| P3, backlog | Medium |
The default, and today the only value the gate issues |
The severity gate as built¶
The whole of the workflow's severity computation is two conditions, in the Compute_Context
step of infrastructure/terraform/modules/datadog-observability/workflows.tf:
const metricsAttrs = (metricsIssues[0] && metricsIssues[0].attributes) || {};
const impactedSessions = metricsAttrs.impacted_sessions ?? 0;
const isRegression = metricsAttrs.is_regression === true || regressionOfKey !== "";
const SESSIONS_THRESHOLD = 10;
const priorityName = isRegression || impactedSessions >= SESSIONS_THRESHOLD ? "High" : "Medium";
| Priority | Mechanical test |
|---|---|
High |
The issue carries a regression flag, or its newest PPL ticket is Done or Closed, or impacted sessions >= 10 |
Medium |
Everything else |
Both inputs come from Get_Issue_Metrics, a searchIssues call for issue.id:<id> windowed
from the issue's first_seen to now. PPL-3451 moved the regression read there on the evidence
that the getIssue payload has no is_regression attribute. The live step output shows the
search returns no such attribute either: a run on 1 October 2026 for an issue with 707
occurrences returned {"total_count": 707} and nothing else. So impactedSessions falls back
to 0, the metrics half of isRegression is false, and the gate answers Medium for every
ticket except a regression on a finished one, where regressionOfKey comes from Jira instead. Any fix has to request
the impact and regression fields explicitly, or read them from a source that returns them, and
then be checked against a real run's step output rather than against the corpus pull.
Two further properties of the window, which matter once the inputs arrive:
- For a newly-seen issue,
first_seento now is close to the 24-hour window the thresholds were derived for. For a regression reopened months later it spans the issue's whole history, so the count is cumulative, but a regression is alreadyHighon the first condition. - The search has no
envfilter, and an Error Tracking issue groups across every environment that reports it, so the counts include preview and staging traffic. PPL-3712 (#2160, open) scopes it toenv:prod.
Every issue gets a ticket and a #tech-errors message regardless of priority, with two
exceptions: the circuit breaker, and a run that fails before its
Slack step. Severity sets the priority and the wording of the Slack message; it never decides
whether the work is recorded.
What the workflow does per issue¶
- Fetch the issue (
getIssue), then fetch its metrics (searchIssues). - Search PPL for an existing ticket by
"Datadog Issue ID" ~ <issue id>, then fail the run if Jira did not authenticate the search. Jira answers an unauthenticated JQL search with200and no results, which would otherwise read as "no ticket" (PPL-3816, see Datadog error triage). - Open ticket found (newest ticket by
created, in a To Do or In Progress status): leave its status alone, add theregressionlabel, comment with the version it returned in, postREGRESSIONto#tech-errors, set the Datadog issue toACKNOWLEDGED. - Finished ticket found (
DoneorClosed): the PPL workflow has no transition from either back to a To Do status, so the ticket cannot be reopened. Take the create path in step 5, with three differences: the newBugisHigh, its description starts withRegression of <old key>, and it gets theregressionlabel and arelates tolink to the old ticket. The Slack post readsREGRESSION of <old key>. Because dedupe takes the newest ticket, the next regression of the same issue updates this new ticket. A failed link falls through to the label, and a failed label to the Slack post. - No ticket found: check the circuit breaker, then create a
Bugin PPL and postNEWto#tech-errors, then set the issue toACKNOWLEDGED.
The ticket the create path files carries:
| Field | Value |
|---|---|
| Summary | [<service>] <error type>: <first line of the message>, cut at the first { and capped at Jira's 254 characters |
| Description | Datadog issue link, service, error type, first-seen and last-seen time and version, total occurrences, impacted sessions, and the full message in a {code} block |
| Priority | High or Medium from the gate |
| Component | From the service to component mapping |
Datadog Issue ID, Error Occurrences, Impacted Sessions, Error Last Seen |
Stamped once, at creation |
| Reporter | The Jira integration account |
A run that fails at any step reaches neither Slack nor Jira past that step, and leaves the
Datadog issue OPEN. The [Workflows] Terraform-managed workflow execution failed monitor is
the only thing that makes that visible.
P1 is parked, and nothing pages anyone¶
There is no P1 tier and no #tech-alerts-p1 channel. PPL-3325 holds that work, and it is
parked on evidence, not on capacity: checked over a 400-day env:prod window, the originally
drafted P1 criteria would have fired zero times.
| Original P1 criterion | Verdict |
|---|---|
is_crash == true |
Dead. No crashes exist. Every issue is platform: BROWSER or BACKEND, and a crash-filtered search returns no data. Meaningful only if native iOS or Android builds start reporting |
| Impacted sessions as a share of active sessions in 15 minutes | Dead. Impact accumulates over weeks rather than spiking; see why the rate expression fails |
| Blocks a clinical flow | Viable, but depends on the criticality reference table, which does not exist |
| New issue on a release less than 2 hours old | Viable and demonstrably matchable, since first_seen_version ties to a real release |
When PPL-3325 is unparked, the direction is to rebase on what actually distinguishes urgency in this data: clinical-flow failure at any volume, a new issue on a fresh release, a rate change against the issue's own baseline, and possibly an absolute count over a longer window (around 50 impacted sessions in 24 hours would have caught exactly the top five issues in 14 months and nothing else). Do not reinstate the crash or 15-minute-share criteria without new evidence that they can match.
Nothing in this pipeline interrupts anybody
The error monitors are notifications at monitor priority 3 and 4. Error messages land in
#tech-errors and workflow failures in #tech-alerts. If an error genuinely needs
someone now, escalate by hand, in the same way as any other unplanned production problem.
Do not assume an alert reached a human because a ticket exists.
High: ticket this sprint¶
Mechanically, either of:
| # | Signal | Test |
|---|---|---|
| 1 | Regression | The issue carries a regression flag, meaning Datadog reopened a RESOLVED issue on a newer version |
| 2 | Sustained session impact | Impacted sessions >= 10 |
Neither reaches the gate today, so apply both by hand at triage: the Datadog issue page shows
the regression and the impacted-session count, which the ticket's Impacted Sessions field
does not.
And by human judgement at triage, promote to High when the error is a repeatable
functional break: a user can reproduce it by following a documented flow. That is the
criterion no workflow can evaluate, and it is the main reason the weekly triage exists.
The derivation below produces a 3 impacted sessions figure as the natural boundary for this tier. The gate does not implement it; it uses a single threshold of 10 for both. Treat 3 as the guide when triaging by hand, and see the gaps table.
Medium: backlog¶
The default. Everything that matched no High criterion, plus anything triage judges to be
cosmetic, single-user, non-blocking, or self-recovering (it degrades then recovers without
user action).
Known gaps between this page and the gate¶
Read this before concluding that the automation is wrong or that this page is stale. Each row is a deliberate divergence or an open defect, not drift.
| Rubric criterion | Gate as built | Why, or what to do |
|---|---|---|
Impacted sessions >= 10 implies High |
Implemented, never fires | searchIssues returns no impacted_sessions, so the value is always 0. Open defect, PPL-4168 |
Regression implies High |
Fires only for a regression on a Done or Closed ticket |
PPL-3451 moved the read to searchIssues, which returns no is_regression either; the finished-ticket case is known from Jira instead. Open defect for the rest, PPL-4168 |
| A regression reopens its ticket | Updates an open ticket, files a new linked Bug for a finished one |
No Done or Closed ticket has a transition back to a To Do status, so a finished ticket is never reopened. PPL-4169 |
| Counts describe prod | Counts span every environment | The metrics search has no env filter. PPL-3712, #2160 (open) |
| Expected 4xx outcomes get no ticket | Some still do | dd-trace marks the middleware span whatever the handler answers. Root cause fixed in PPL-3661; PPL-3712 adds a monitor exclusion as a backstop |
is_crash == true implies P1 |
Not implemented | No crash has ever been recorded. Parked with PPL-3325 |
| Clinical-flow failure implies P1 | Not implemented | Needs the criticality table, which needs Product sign-off |
| New issue on a release under 2 hours old implies P1 | Not implemented | Viable, parked with PPL-3325. Needs a deploy-timestamp lookup, which is not on the issue object |
| Session share over 15 minutes | Not implemented | Not viable at our traffic, and the monitor uses new() with no rate term at all |
High at 3 impacted sessions in 24 h |
Single threshold of 10, over first_seen to now |
Worth reconciling once the gate can see sessions at all: either lower the constant to 3, or restate the tier at 10 and drop the 3 |
| Every new issue gets a ticket | True except under the circuit breaker or a failed run | By design for the breaker, see issues that get no ticket |
How the thresholds were chosen¶
Pulled 5 August 2026 over a trailing 30-day window (env:prod, so roughly 6 July to
5 August 2026), paging the whole corpus rather than sampling the top of the list.
Corpus shape. 150 prod Error Tracking issues, which reconciles exactly:
| Service | Issues | Carry impacted_sessions |
Events |
|---|---|---|---|
clinicians-app-frontend |
50 | 50 | 35,675 |
members-app-frontend |
40 | 40 | 6,071 |
perci-platform-backend |
60 | 0 | 14,800 |
The 60 backend issues carry no impacted_sessions attribute at all: they are
platform: BACKEND, which has no session concept. Any session-based threshold is
therefore frontend-only, and with the crash, criticality and fresh-release criteria all
parked, backend severity rests on the regression flag alone until PPL-3325 is picked up.
The percentiles below are over the 90 frontend issues that have the attribute, because
computing a percentile over 150 values where 60 are absent would be arithmetic on a hole.
Distribution of impacted_sessions across the 90 frontend issues (30 days):
| Statistic | Value |
|---|---|
| min / median / max | 1 / 2 / 280 |
| mean | 16.0 |
| p50 / p75 / p90 | 2 / 8 / 18 |
| p95 / p99 | 105 / 280 |
The distribution is extremely top-heavy: the 5 largest issues hold 67.9 % of all impacted sessions, and the top 3 hold 50.1 %. The upper tail has a clean break, with values running 18, 20, 39, then jumping to 57, 105, 151, 165, 276, 280.
Converting to a per-day rate. impacted_sessions is measured over whatever window you
query, so a 30-day figure is cumulative and cannot be compared to a 24-hour threshold
directly. Dividing by the 30-day window:
| Issue rank | 30-day impacted sessions | Mean per day |
|---|---|---|
| 1 | 280 | 9.3 |
| 2 | 276 | 9.2 |
| 3 | 165 | 5.5 |
| 4 | 151 | 5.0 |
| 5 | 105 | 3.5 |
| 6 | 57 | 1.9 |
This gives both numbers the epic asked for:
- 10 impacted sessions in 24 hours as the urgent threshold. No issue in the corpus
averages more than 9.3 per day, so nothing then in prod would have tripped it on volume
alone. That was the intent: the corpus contained no acute outbreak, only slow burns
accumulated over a month. On an app serving roughly 70 to 120 sessions a day, 10 affected
sessions in one day is a blast radius of order 10 %. With P1 parked, this is the constant
the gate now uses for
High. - 3 impacted sessions in 24 hours as the boundary for ordinary-but-real impact. Exceeded by 5 of 90 issues on a per-day basis, so about the 94th percentile of observed per-day rates: high enough that the median issue (0.07 sessions per day) never reaches it, low enough to catch a real functional break early. The gate does not use it; triage should.
Why 24 hours was the natural window, and what the workflow actually does
The Datadog issue-detail endpoint takes no time parameters: its figures always cover the
24 hours before the issue was last seen, so a 24-hour count comes free with fetching the
issue. The shipped workflow instead runs a searchIssues call from the issue's
first_seen to now, which for a new issue is close to the same number and for an old one
is cumulative. Both thresholds above are stated against 24 hours because that is the
window they were derived for. As called today, that search returns no
impacted_sessions at all, so neither threshold is evaluated (see
the severity gate as built).
The rate expression is not viable at current traffic¶
The epic asked for two expressions of the session criterion: a rate against active sessions for the monitor, and an absolute count for the workflow. The absolute count is above. The rate, as originally specified over a 15-minute window, does not work at our traffic level, and the number matters more than the shape of the rule.
Prod RUM sessions (@type:session, @session.type:user) on 21 July 2026, the busiest
day in the retained window, in 15-minute buckets:
| App | Sessions that day | Non-empty 15-min buckets | Peak bucket | Median non-empty bucket |
|---|---|---|---|---|
| Member App | 122 | 53 of 96 | 6 | 2 |
| Clinical App | 71 | 44 of 96 | 4 | 1 |
At a peak of 6 concurrent sessions per 15 minutes, one session is 16.7 % of the busiest window on the busiest day, and 50 % of a median window. Any percentage threshold is therefore decided by a single user: 10 % of the peak bucket is 0.6 sessions, so the threshold rounds down to "one error". A rule like that looks quantitative and discriminates nothing.
This is why PPL-3273 shipped a monitor that counts new issues rather than a share of sessions, and why the share-based P1 criterion is parked rather than re-tuned. The best rate expression available at this traffic level, if one is ever wanted, is a 6-hour window at 20 % of active sessions with a floor of 5 impacted sessions: at roughly 5 sessions an hour on a weekday, 6 hours accumulates 25 to 30 sessions per app, the smallest denominator for which 20 % still requires 5 distinct impacted sessions and cannot be tripped by one user. The floor is what keeps the quiet hours safe.
When to revisit the 15-minute window. For a 20 % threshold to require 5 impacted sessions within 15 minutes, a single app needs about 25 active sessions per 15 minutes, which is roughly 100 an hour or 2,400 a day. Current peak is 6 per 15 minutes. So a 15-minute rate becomes meaningful at roughly 20 times today's traffic. Re-run the bucket query above before adopting it.
Clinical criticality (proposal)¶
Not built, and not signed off
The flow list below is a proposal from engineering, seeded from the flows named in PPL-3271 plus the view paths and backend modules that actually carry prod errors. Product has not agreed it, no reference table exists, and PPL-3274 shipped without one. It is now an explicit unpark condition for PPL-3325: clinical-flow detection is the criterion that matters most for a health platform and the one with the least data behind it.
Service-level tiering is useless here: there are three services and all three matter. The signal has to be at flow level, so criticality belongs in a Datadog Reference Table, keyed on the service plus the code location, so Product can change what counts as critical by editing data rather than by editing a workflow.
Intended table shape:
| Column | Type | Meaning |
|---|---|---|
service |
string | members-app-frontend, clinicians-app-frontend, perci-platform-backend |
match_type |
string | view_path, file_path, or function_name |
match_value |
string | The value to match, exactly |
criticality |
string | clinical-critical or standard |
flow |
string | Human-readable flow name, used in the ticket summary and the Slack message |
Primary key is (service, match_type, match_value). A lookup that misses returns
standard, so the table is a positive allowlist and an unmapped path never silently
becomes urgent.
Proposed clinical-critical rows. The view paths are the real paths carrying prod
errors today; the backend files are the real modules that produced prod issues.
| Flow | Service | match_type |
match_value |
|---|---|---|---|
| Joining a call | clinicians-app-frontend |
view_path |
/appointmentCall |
| Joining a call | members-app-frontend |
view_path |
/videoAppointment |
| Joining a call | perci-platform-backend |
file_path |
src/bff-clinical/appointments/endAppointmentCall.ts |
| Appointment conclusion | clinicians-app-frontend |
view_path |
/appointmentConclusion |
| Appointment conclusion | clinicians-app-frontend |
view_path |
/appointmentConclusionSummary |
| Appointment conclusion | perci-platform-backend |
file_path |
src/bff-clinical/appointments/upsertAppointmentConclusion.ts |
| Referral intake | members-app-frontend |
view_path |
/referralSignup |
| Referral intake | perci-platform-backend |
file_path |
src/bff/referrals/verifyReferralTokenData.ts |
| Login | members-app-frontend |
view_path |
/login |
| Login | members-app-frontend |
view_path |
/start-sso |
| Login | members-app-frontend |
view_path |
/sso-callback |
| Login | clinicians-app-frontend |
view_path |
/signin |
| Login | perci-platform-backend |
file_path |
src/bff/middleware/descopeMemberMiddleware.ts |
| Login | perci-platform-backend |
file_path |
src/bff-clinical/middleware/descopeClinicianMiddleware.ts |
Candidates engineering suggests Product also considers, because they carry prod errors
and plausibly block care rather than merely inconveniencing a user: appointment booking
(/bookAnAppointment, src/bff/appointments/createAppointment.ts), the code and SSO
signup chain (/ssoSignupHome, /codeSignupAbout, /codeSignupBasicInfo), and clinical
messaging (/MessagesHome, /newMessagesHome). These are not in the proposed list
above; they are the agenda for the Product conversation.
Backend file_path values need normalising
Backend issues report file_path with the container prefix, for example
/workspace/src/bff/referrals/verifyReferralTokenData.ts. Strip the /workspace/
prefix before the lookup, or store the prefixed form. Note also that 16 of the 60
backend issues report a file_path inside node_modules (13 of them in
@grpc/grpc-js), which no criticality row will ever match. Those fall through to
standard correctly.
Datadog states¶
Five states exist in the public API. These are the only values
PUT /api/v2/error-tracking/issues/{issue_id}/state accepts:
| State | Meaning here |
|---|---|
OPEN |
Never triaged. Every issue the workflow handles successfully is moved to ACKNOWLEDGED, so an OPEN issue that has alerted means the workflow run failed. Today that is mostly regressions, see the live defects |
ACKNOWLEDGED |
Recorded and being worked, or suppressed by the circuit breaker |
RESOLVED |
Fixed and released to production. Reopens automatically if it recurs on a newer version |
IGNORED |
Deliberately never actioned. Stays silent forever, including on new versions |
EXCLUDED |
Not a real error, should never have been recorded. Use this for taxonomy and environment-tagging mistakes, not for "we chose not to fix it" |
FOR_REVIEW and REVIEWED are not API values
Some tooling, including the Datadog MCP server and parts of the UI, uses FOR_REVIEW
and REVIEWED as aliases for OPEN and ACKNOWLEDGED. They are aliases only.
Sending either to the state endpoint fails. When reading state from a tool that
returns them, map FOR_REVIEW to OPEN and REVIEWED to ACKNOWLEDGED before
comparing against anything on this page.
Why RESOLVED and IGNORED are not interchangeable¶
This distinction is the entire basis of regression detection, and it is the one state decision worth getting right:
- A
RESOLVEDissue that recurs on a newer version reopens itself and is flagged as a regression. That is free regression detection, with no memory and no spreadsheet. - An
IGNOREDissue never comes back, on any version, ever.
So IGNORED is not "closed"; it is "permanently deaf to this error". Reach for it only
when the error is genuinely acceptable forever. If there is any chance the fix matters
later, RESOLVED is the safer choice: the worst case is one reopened ticket, whereas a
wrongly ignored issue is a silent defect with no alarm attached.
Datadog also resolves issues on its own, too early
Datadog's GitHub integration sets an issue RESOLVED the moment a PR it has linked to
the issue merges into develop, days before the fix reaches production, and whether or
not the PR fixed anything. An occurrence in that gap reopens the issue as a false
regression. The fix is a Datadog console change (turn off PR-based auto-resolve), tracked
on PPL-3711. Until then, a RESOLVED issue does not prove the fix is released.
For scale, the 5 August corpus was 98 OPEN, 12 ACKNOWLEDGED, 37 RESOLVED and 3
IGNORED, and 59 of the 150 issues already carried a regression flag. State then was close
to meaningless, which is the problem the pipeline exists to fix.
Issues that get no ticket¶
The workflow carries a circuit breaker, and it is the one designed case where an issue is
ACKNOWLEDGED with no ticket behind it. Before creating a ticket it counts PPL tickets with
a non-empty Datadog Issue ID created in the last 15 minutes; if there are more than 5,
it posts a Circuit breaker tripped message to #tech-errors, sets the issue to
ACKNOWLEDGED, and creates nothing.
That protects the board from an error storm at the cost of losing the tickets for it. So a
circuit-breaker message in #tech-errors is itself a triage action: work out what broke,
raise one ticket for it by hand, and check Error Tracking for ACKNOWLEDGED issues from that
window with no Datadog Issue ID recorded against them.
The other case is undesigned: a run that fails. It leaves the issue OPEN with no ticket,
or with an untouched ticket on the regression path, and announces itself only through
[Workflows] Terraform-managed workflow execution failed in #tech-alerts.
Ticket status to Datadog state¶
Three Jira automation rules on project PPL (PPL-3275), all gated on Datadog Issue ID being
non-empty, each a Send web request to
PUT https://api.datadoghq.eu/api/v2/error-tracking/issues/{issue_id}/state with a
service-account app key scoped to error_tracking_read and error_tracking_write only.
Done means released to production in this project (see
Delivery), which is why Done maps to RESOLVED rather
than to something earlier. The release chain gets there on its own: a release PR merging to
main moves its tickets from Ready for Release to Done, and the Done rule fires.
| Rule (as named in Jira) | Trigger | Datadog state |
|---|---|---|
| Set by the workflow, not a rule | Ticket created or updated by the workflow | ACKNOWLEDGED |
Datadog: Done -> RESOLVED |
Transition to Done |
RESOLVED |
Datadog: Closed as Won't Do or Declined -> IGNORED |
Transition to Closed, and resolution in ("Won't Do", "Declined") |
IGNORED |
| (no rule) | Transition to Closed with any other resolution |
Left as it was, normally ACKNOWLEDGED |
Datadog: reopened out of Done -> ACKNOWLEDGED |
Transition out of Done to any status |
ACKNOWLEDGED |
Each rule has "delay execution until web request response" enabled, so a failure shows up in the automation audit log instead of dropping silently. Rule-editing on PPL is restricted, because Jira Automation has no secret store: anything in a web-request header is readable by anyone who can edit the rule.
Which resolutions count as won't-do¶
Project PPL had exactly six resolution values when this was checked. Counts are live usage
across the project as of 5 August 2026, and confirm the list was complete
(529 + 55 + 47 + 24 + 22 + 1 = 678, which equalled project = PPL AND resolution IS NOT EMPTY).
| Resolution | Issues | Should map to IGNORED? |
Live rule |
|---|---|---|---|
| Done | 529 | No. Fixed, so RESOLVED |
Not IGNORED |
| Duplicate | 55 | No. See below | Not IGNORED |
| Won't Do | 47 | Yes | IGNORED |
| Cannot Reproduce | 24 | No. Unexplained, not accepted | Not IGNORED |
| Functioned As Designed | 22 | Yes | Not IGNORED. The rule omits it |
| Declined | 1 | Yes | IGNORED |
Won't Do, Functioned As Designed and Declined are the won't-do style decisions.
The live rule covers only Won't Do and Declined, so an error closed as Functioned As
Designed stays ACKNOWLEDGED and alerts again only as a regression. Add it to the rule's JQL
condition, or change this table if leaving it out is intended.
A duplicate must never set IGNORED
Closing a ticket as Duplicate says nothing about whether the error is being fixed,
only that this ticket is not where the fix lives. Setting IGNORED would silence an
error that another open ticket is still working on, and it would silence it
permanently. Leave a duplicate at ACKNOWLEDGED and let the surviving ticket drive the
state. Same reasoning for Cannot Reproduce: an error we could not reproduce is
unexplained, not accepted, and it should be allowed to come back and tell us more.
The resolution may arrive after the transition
In PPL the Closed and No testing needed transitions carried no resolution screen
when checked: getTransitionsForJiraIssue returned empty fields for every transition.
Resolutions were set as a separate field edit afterwards. In PPL-2929, for instance, the
status moved to Done and the resolution was set two seconds later as its own changelog
entry. The IGNORED rule is triggered by the transition and evaluates its resolution JQL
at that moment, so if the resolution lands afterwards the condition fails and the rule
does nothing. Set the resolution in the same action as the close, or check the automation
audit log after closing. No auto-filed ticket has yet been closed as Won't Do or Declined,
so the rule has not been exercised in production. Adding resolution screens to those
transitions would be the clean fix.
Service to component mapping¶
Tickets the workflow creates carry the Component matching the Datadog service, so they route on the PPL board like any other ticket (see Delivery).
Datadog service |
PPL Component | Component id |
|---|---|---|
clinicians-app-frontend |
Clinical App | 10075 |
members-app-frontend |
Member App | 10074 |
perci-platform-backend |
Backend | 10330 |
All three are mapped explicitly. The workflow first shipped with only
clinicians-app-frontend in the map, so member app errors filed under Backend until PPL-3450
added the Member App row. An unmapped service still defaults to Backend on purpose, so a
service nobody has mapped yet lands somewhere a human sees it. web-professionals is one
such service today: it reports to prod Error Tracking and files under Backend (for example
PPL-4035).
The Datadog Issue ID field¶
The join key between a Datadog issue and its PPL ticket, added by PPL-3270.
| Property | Value |
|---|---|
| Field id | customfield_10798 |
| Name | Datadog Issue ID |
| Type | Text Field (single line) |
| Scope | Project PPL, optional, on the Bug, Story and Task screens |
| JQL | "Datadog Issue ID" ~ "<uuid>" |
| Saved filter | project = PPL AND "Datadog Issue ID" IS NOT EMPTY |
The workflow stamps it on creation and uses it three ways: to find the existing ticket when an issue reopens rather than creating a second one, to count recent auto-created tickets for the circuit breaker, and to scope the PPL-3275 automation rules so they never fire on a hand-written ticket.
On 2 October 2026, 172 PPL tickets carried it: 156 filed by the workflow and 16 stamped by
hand. Ten Datadog issue ids appear on more than one ticket. Seven are pairs of hand-stamped
tickets from before the workflow. The other three are workflow duplicates filed while the
dedupe search was unauthenticated (PPL-3816, fixed 15 September): one issue five times in
August, and two issues twice each in early September. Dedupe asks Jira for the newest match
(ORDER BY created DESC), so for those ids the regression path acts on the latest ticket.
Frequency and impact on the ticket¶
PPL-3352 added three custom fields that the workflow stamps when it creates a ticket:
| Field | Id | Source |
|---|---|---|
Error Occurrences |
customfield_10909 |
total_count from the metrics search |
Impacted Sessions |
customfield_10910 |
impacted_sessions from the metrics search, so always 0 today (see the defect) |
Error Last Seen |
customfield_10911 |
The issue's last_seen |
They are a creation-time snapshot. Nothing refreshes them: the triage workflow is the only Datadog workflow, and it writes these fields only on the create path. An error that arrives as 12 occurrences and grows to 4,000 still reads 12 on its ticket months later, which is the difference between a backlog you can prioritise and a pile of equally-sized cards.
The intended end state is a scheduled refresh over open error-derived tickets, so the triage
queue becomes a saved filter sorted by Impacted Sessions descending, and a ticket triaged
Medium months ago that has quietly become the highest-impact issue on the board becomes
visible without anyone remembering to re-check it. That needs both the refresh job and a
metrics search that returns impacted_sessions. Until then, read impact off the Datadog issue
page, not the ticket.
The weekly triage¶
A half-hour slot, one owner per week, rotating round-robin through the engineers on the
@percihealth/backend and @percihealth/frontend teams (the same two teams that drive
PR review routing). PPL-3276 starts the
rotation and retires the manual path. It has not started.
The owner works one saved filter, not the Datadog UI. New auto-filed tickets land in Triage:
- Set or correct the priority. Every ticket the workflow created since the last session.
While the gate is blind every one arrives
Medium, so this step is where allHighdecisions are made: open the Datadog issue link, apply theHighcriteria by hand, and add the judgement criteria (repeatable functional break, cosmetic) that the gate can never evaluate. Until the criticality table exists, it is also where a clinical-flow failure gets the priority no automation gave it. - Clear anything untriaged. Nothing may stay untriaged for more than seven days. If a ticket has had no priority confirmed within seven days of creation, it is the triage owner's job to give it one that week, or to close it with a resolution.
- Pick up what the pipeline dropped. Any
Circuit breaker trippedmessage in#tech-errors, and any[Workflows] Terraform-managed workflow execution failedalert in#tech-alerts, since the last session. Both mean issues exist with no ticket, or with a regression nobody was told about. See issues that get no ticket. - Feed the corrections back. A priority that was wrong for a systematic reason is a change to this page or to the gate, not a one-off edit. Raise it.
CONFIRM: rota and slot
The named rota and the day still need confirming with the team, alongside the wider on-call gaps flagged in Team ownership. There is no out-of-hours escalation path to define yet, because nothing pages: see P1 is parked.
Engineers do not open Datadog¶
Deliberately. The whole pipeline exists so that Error Tracking state is derived, not maintained by hand:
- Errors arrive as PPL tickets, and are worked in Jira and GitHub like any other ticket.
- Datadog state follows the ticket status automatically, per the mapping above.
- Nobody is expected to log in to mark something reviewed or fixed. That expectation is precisely why Error Tracking state was meaningless before the pipeline.
Open Datadog to investigate an error (stack trace, affected versions, session replay, and today its impacted-session count), never to record a decision. If you find yourself changing state by hand, either the ticket is missing or an automation is broken, and both are worth a ticket of their own.
Related¶
- Error taxonomy: expected client outcomes versus defects.
- Datadog error triage: environment reporting gates, how the monitor reaches the workflow, the workflow's Jira authentication, and how to read an Error Tracking issue.
- Datadog backend observability: how backend telemetry is produced.
- Delivery: the Jira statuses and resolutions this page maps from.