OPERATIONS
Platform Health and Site Reliability Engineering
Know what failed, take ownership, and verify recovery.
SERVICE RELIABILITY
Are we meeting our targets?
LIVE LEGACY OPERATIONS
Last 15 minutes
Active alarm inventory
Aggregate AWS infrastructure metrics only. No request paths, logs, clinical payloads, or patient identifiers are queried.
Monthly SLA tracking
Configure SLOs and SLA tracking
Draft internal targets. Saving recalculates the displayed periods using these targets; it does not approve a customer agreement or alter historical measurements.
SLIs are good events ÷ total events. Error budget = total events × (1 − target). Burn rate compares measured error rate to the allowed error rate. Missing hours are unknown, not healthy. Readiness is probe-based availability, not a guarantee of every user workflow. Burn warnings appear here; they do not yet send SMS or block deployments.
NEXT ACTION
Inspect current health and the recent deployment. Keep patient information out of incident notes and messages. Verify uncertain clinical writes before retrying.
Activity
INCIDENT
Next action
Actions update this incident and its activity history.
Operational access only
Use your organization sign-in to view incidents. An operator role is required to acknowledge or snooze alerts.