Skip to content

Investigating incidents

Connect the telemetry service, source repository, deployment system, and data source for the incident.

Investigate checkout errors from 09:00 to 10:00 Pacific. Correlate Datadog logs and traces with deployments and source changes. Report observed facts, likely causes, and uncertainty. Do not change production.

State the environment, service, time range, symptom, and affected users. Include available request identifiers, error text, and monitor links.

Ask the agent to separate:

  • Current observed state.
  • Events that happened before or during the incident.
  • Causal evidence.
  • Inference and unresolved questions.

An application error does not prove that a database, provider, or network is unhealthy. Confirm each dependency from its own evidence.

Read-only investigation is the default. A restart, rollback, configuration change, or data repair needs a separate request. State the exact target and blast radius. Ask for verification after an approved operation.

See the complete incident investigation recipe.