Site Reliability Engineer
role · reused
Evidence record
- What is verified
- Uses in published workflows: 1
- Revision
- Version 0.1.5
- Validation status
- Public-sanitized publication; no independent validation is claimed.
What this is and why
The role that asks whether a system will survive production. It defines service-level objectives and error budgets, designs observability (metrics, structured logging, tracing) and an alerting strategy that pages only on what demands action, and plans incident response before the incident. Its habit of mind is chaos engineering: assume the dependency will be slow and the service will die, and make degradation, rollback, and recovery designed behaviors. Failures end in blameless postmortems with actionable follow-ups.
Metadata and provenance
- id
- datarim-role-sre
- type
- role
- version
- 0.1.5
- origin
- created in Arcanada
- lifecycle
- public_sanitized
- tags
- role, talo-0029
The source is the private knowledge repository; only the sanitized public projection reaches this site.
verified · datarim-role-sre · v0.1.5