Security and reliability are operating habits, not launch-day stickers.
This playbook explains the baseline questions we use to make ownership, access, change and recovery visible. Controls are adapted to the system and contract; this is not a certification claim.
01
Know the assets and boundaries
Record the data classes, environments, integrations, administrators and external services in scope. A diagram must show where trust changes—not merely where servers sit. Unknown scheduled jobs and inherited credentials become explicit findings.
Named information and system owners
Data flow and third-party register
Environment and integration inventory
02
Make access deliberate
Roles follow actual job tasks and high-impact actions receive narrower privileges. Joiner, mover and leaver paths are documented. Service accounts, API keys and emergency access have owners, storage rules and review dates.
Least-privilege role matrix
Multi-factor authentication where supported
Privileged and dormant account review
03
Release with evidence
A release record links the approved change, review evidence, test result, known limitations and rollback decision. Urgent fixes may use an expedited path, but the retrospective evidence still has to be completed.
Peer review and dependency visibility
Environment-specific configuration checks
Approval and rollback ownership
04
Prepare for imperfect days
Alerts should point to an action. Incident contacts, severity definitions and communication responsibilities are agreed before a crisis. Backups count as a control only when restoration has been tested at an appropriate interval.
Useful logs with proportionate retention
Incident triage and communication guide
Restore or recovery exercise evidence
05
Close findings visibly
Findings include consequence, evidence, owner and target. Closure means the control is implemented and checked—not that a ticket was moved. Accepted risk is time-bound and approved by someone with authority.