Skip to content

Datadog Backend Observability (Firebase Functions Gen2)

This document defines backend Datadog setup for: - Logs: Cloud Logging -> Pub/Sub -> Dataflow -> Datadog - Traces: post-deploy Cloud Run instrumentation via datadog-ci

Infrastructure (Terraform)

Terraform creates the GCP + Datadog log pipeline in: - infrastructure/terraform/modules/datadog-observability - infrastructure/terraform/environments/staging - infrastructure/terraform/environments/production

Required Terraform workspace variables: - datadog_logs_api_key (sensitive)

Required Terraform workspace environment variables for Datadog provider: - DATADOG_API_KEY - DATADOG_APP_KEY

CI/CD Instrumentation

The deploy-backend job of .github/workflows/deploy.yml runs Datadog instrumentation after the Firebase Functions deploy, for both environments. The command used is: - npx -y @datadog/datadog-ci@5.21.2 cloud-run instrument ... --tracing true

Current services instrumented: - api - bff - bff-clinical - bff-clinical-media

If a new HTTP function is added, add its service name to that step. Two neighbouring lists in the same job need the same addition (static IP egress and deployment verification). See Adding a component.

Transient Dependency Failures

An error is recorded only when a retry budget is exhausted and the operation actually failed. A dependency failure that the retry layer absorbs is a warning plus a counter. This is configured in apps/perci-platform-backend/functions/src/observability/datadogTracing.ts:

  • DD_GRPC_CLIENT_ERROR_STATUSES excludes the gRPC status codes the Google client libraries retry internally — 4 DEADLINE_EXCEEDED, 10 ABORTED, 14 UNAVAILABLE. Without this, normal Firestore transaction contention (10 ABORTED) dominates Error Tracking even though the client succeeds on retry. dd-trace reads this only from the environment, and only while tracer.init builds its config, so it is set in code immediately before that call.
  • The dns plugin is disabled. Its spans only cover the client libraries' own SRV/TXT lookups, where a host with no TXT record answers ENODATA — a normal negative answer, not a failure.
  • Cloud Logging write failures in Logger are reported with console.warn, not console.error: a failed log shipment must not be raised as a production error through the pipeline it just failed to reach.

Retry volume stays observable:

  • gRPC: dd-trace still tags each span with rpc.grpc.status_code, so contention is queryable and alertable in APM independently of Error Tracking. For Firestore contention, alert on the rate of service:perci-platform-backend @rpc.grpc.status_code:10 spans, not on error count.
  • HTTP: fetchWithRetry emits the perci.backend.fetch.retry counter, tagged outcome:absorbed|exhausted|not_retried and target:<host>. Exhausted attempts are logged at ERROR with the attempt count, and FetchRetryError carries attempts.

Verification Checklist

Run after staging deploy, then production deploy.

  1. Confirm Dataflow is healthy
  2. GCP Console -> Dataflow -> job datadog-export-job-staging / datadog-export-job-prod
  3. Status should be Running with no sustained errors.

  4. Confirm logs arrive in Datadog

  5. Query in Datadog Logs:
  6. service:perci-platform-backend env:staging
  7. service:perci-platform-backend env:production

  8. Confirm traces arrive in Datadog APM

  9. Service catalog should show perci-platform-backend.
  10. Check recent traces for endpoints under api, bff, bff_clinical, bff_clinical_media.

  11. Confirm log/trace correlation

  12. Open a trace span and verify related logs are linked.
  13. Open a log entry and verify dd.trace_id and dd.span_id exist in payload.

Rollout Gates

Use this sequence for safe rollout:

  1. Apply Terraform in staging.
  2. Deploy staging backend and verify checklist above.
  3. Keep staging stable for at least one deploy cycle.
  4. Apply Terraform in production.
  5. Deploy production backend and verify checklist above.

Failure Recovery

If instrumentation fails in CI: 1. Re-run backend deploy workflow. 2. Run manual command with dry-run first: - npx -y @datadog/datadog-ci@5.21.2 cloud-run instrument --project <project> --region europe-west2 --service <service> --tracing true --env <env> --version <sha> --dry-run 3. If needed, temporarily disable Datadog instrumentation step and proceed with deploy, then remediate in a follow-up run.

If log forwarding fails: 1. Check Dataflow worker logs. 2. Validate Secret Manager access for Dataflow worker SA. 3. Validate sink writer has roles/pubsub.publisher on datadog-export-topic.