Datadog Backend Observability (Firebase Functions Gen2)¶
This document defines backend Datadog setup for:
- Logs: Cloud Logging -> Pub/Sub -> Dataflow -> Datadog
- Traces: post-deploy Cloud Run instrumentation via datadog-ci
Infrastructure (Terraform)¶
Terraform creates the GCP + Datadog log pipeline in:
- infrastructure/terraform/modules/datadog-observability
- infrastructure/terraform/environments/staging
- infrastructure/terraform/environments/production
Required Terraform workspace variables:
- datadog_logs_api_key (sensitive)
Required Terraform workspace environment variables for Datadog provider:
- DATADOG_API_KEY
- DATADOG_APP_KEY
CI/CD Instrumentation¶
The deploy-backend job of .github/workflows/deploy.yml runs Datadog instrumentation after
the Firebase Functions deploy, for both environments. The command used is:
- npx -y @datadog/datadog-ci@5.21.2 cloud-run instrument ... --tracing true
Current services instrumented:
- api
- bff
- bff-clinical
- bff-clinical-media
If a new HTTP function is added, add its service name to that step. Two neighbouring lists in the same job need the same addition (static IP egress and deployment verification). See Adding a component.
Transient Dependency Failures¶
An error is recorded only when a retry budget is exhausted and the operation actually failed.
A dependency failure that the retry layer absorbs is a warning plus a counter. This is configured
in apps/perci-platform-backend/functions/src/observability/datadogTracing.ts:
DD_GRPC_CLIENT_ERROR_STATUSESexcludes the gRPC status codes the Google client libraries retry internally —4 DEADLINE_EXCEEDED,10 ABORTED,14 UNAVAILABLE. Without this, normal Firestore transaction contention (10 ABORTED) dominates Error Tracking even though the client succeeds on retry. dd-trace reads this only from the environment, and only whiletracer.initbuilds its config, so it is set in code immediately before that call.- The
dnsplugin is disabled. Its spans only cover the client libraries' own SRV/TXT lookups, where a host with no TXT record answersENODATA— a normal negative answer, not a failure. - Cloud Logging write failures in
Loggerare reported withconsole.warn, notconsole.error: a failed log shipment must not be raised as a production error through the pipeline it just failed to reach.
Retry volume stays observable:
- gRPC: dd-trace still tags each span with
rpc.grpc.status_code, so contention is queryable and alertable in APM independently of Error Tracking. For Firestore contention, alert on the rate ofservice:perci-platform-backend @rpc.grpc.status_code:10spans, not on error count. - HTTP:
fetchWithRetryemits theperci.backend.fetch.retrycounter, taggedoutcome:absorbed|exhausted|not_retriedandtarget:<host>. Exhausted attempts are logged atERRORwith the attempt count, andFetchRetryErrorcarriesattempts.
Verification Checklist¶
Run after staging deploy, then production deploy.
- Confirm Dataflow is healthy
- GCP Console -> Dataflow -> job
datadog-export-job-staging/datadog-export-job-prod -
Status should be
Runningwith no sustained errors. -
Confirm logs arrive in Datadog
- Query in Datadog Logs:
service:perci-platform-backend env:staging-
service:perci-platform-backend env:production -
Confirm traces arrive in Datadog APM
- Service catalog should show
perci-platform-backend. -
Check recent traces for endpoints under
api,bff,bff_clinical,bff_clinical_media. -
Confirm log/trace correlation
- Open a trace span and verify related logs are linked.
- Open a log entry and verify
dd.trace_idanddd.span_idexist in payload.
Rollout Gates¶
Use this sequence for safe rollout:
- Apply Terraform in staging.
- Deploy staging backend and verify checklist above.
- Keep staging stable for at least one deploy cycle.
- Apply Terraform in production.
- Deploy production backend and verify checklist above.
Failure Recovery¶
If instrumentation fails in CI:
1. Re-run backend deploy workflow.
2. Run manual command with dry-run first:
- npx -y @datadog/datadog-ci@5.21.2 cloud-run instrument --project <project> --region europe-west2 --service <service> --tracing true --env <env> --version <sha> --dry-run
3. If needed, temporarily disable Datadog instrumentation step and proceed with deploy, then remediate in a follow-up run.
If log forwarding fails:
1. Check Dataflow worker logs.
2. Validate Secret Manager access for Dataflow worker SA.
3. Validate sink writer has roles/pubsub.publisher on datadog-export-topic.