Appearance
Emergency Mode Operation Plan
Requirement: HIPAA §164.308(a)(7)(ii)(C) — Emergency Mode Operation Plan. Establish and implement procedures to enable continuation of critical business processes for protection of the security of ePHI while operating in emergency mode.
Owner: Rory (Security Officer) Audience: Rory, Kevin, Greg, Adriana Activation: This plan is activated when the primary Azure environment is unavailable or significantly degraded for an extended period (>2 hours for critical systems).
This document answers: during an outage, what continues, what pauses, who decides, and what are the minimum acceptable security controls while degraded?
For the technical DR failover procedure (how to start AWS systems), see DR Failover Procedure.
Business Processes and Criticality
Tier 1 — Critical (must continue or restore within 4 hours)
| Process | System Dependency | PHI Involved | Emergency Alternative |
|---|---|---|---|
| Medicare patient access to DDE AVD | DDE tenant (Azure DDE) | Yes — primary PHI access plane | DR failover to AWS (see §2); or notify CMS/customers of outage and document |
| Break-glass account availability | All 18 break-glass accounts | No (access management) | Credentials in Keeper — accessible independently of Azure |
| Incident detection and response | SecOps platform + monitoring agents | Indirect | Arctic Wolf MDR continues independently — they have direct Azure/AWS access and will detect and alert regardless of SecOps platform status |
| Security logging | Log Analytics + CloudTrail | No (logs) | CloudTrail continues independently in AWS; Azure Monitor continues if Azure control plane is up; archive storage has 6-year retention |
Tier 2 — Important (restore within 24 hours)
| Process | System Dependency | PHI Involved | Emergency Alternative |
|---|---|---|---|
| Internal staff email and communications | M365 (Exchange Online) | May contain PHI references | M365 is independent of Azure infrastructure — continues unless Microsoft outage |
| Internal IT operations | Azure PROD VMs | No | Access via break-glass accounts; critical tickets in SecOps or by phone |
| Automated security monitoring (agents) | Container Apps Jobs | No | Arctic Wolf continues; agents resume on Container App restart; email alerts from any ops-automation workflows that remain up |
| Backup replication (Veeam) | Veeam → AWS S3 | No (backup metadata) | Next nightly run after restoration; accept RPO extension during outage |
Tier 3 — Deferrable (restore within 72 hours)
| Process | System Dependency | PHI Involved | Emergency Alternative |
|---|---|---|---|
| Infrastructure changes (Terraform) | GitHub Actions + Azure | No | All changes paused during emergency — freeze applied |
| Compliance reporting | ops-automation pipelines | No | Manual reporting; defer non-urgent reports |
| Vulnerability scanning | GitHub Actions (monthly) | No | Defer to next scheduled run |
| Cost reporting | ops-automation | No | Defer |
Emergency Activation Decision
Who can declare emergency mode: Rory (primary). Kevin or Greg (if Rory is unreachable).
Activation thresholds:
| Condition | Response |
|---|---|
| DDE AVD unavailable > 2 hours with no clear restoration ETA | Activate Tier 1 emergency procedures — consider DR failover |
| Azure PROD unavailable > 4 hours | Activate Tier 2 procedures; evaluate Tier 1 if DDE is also affected |
| Both Azure regions unavailable | Activate full emergency mode; initiate DR failover |
| Active ransomware or security incident | Activate regardless of duration — follow Incident Response concurrently |
Emergency Mode Procedures
1. DDE AVD Unavailable (Medicare Patients Cannot Access Application)
DDE is the highest priority. If Medicare patients cannot reach the published app:
Immediate actions (0–30 minutes):
- Verify the outage — attempt login to DDE AVD portal; check Azure Portal for DDE tenant status
- Check Azure Service Health for known regional outages
- Notify Adriana — she holds the customer communication contacts and CMS relationship
- Post internal status notice via M365 Teams
If restoration expected within 4 hours:
- Wait and monitor; inform customers via Adriana of delay
- No DR failover unless instructed by Rory
If restoration not expected within 4 hours:
- Rory evaluates DR failover to AWS per DR Failover Procedure
- Adriana notifies Medicare customers and CMS of outage and estimated restoration
- Document the outage start time, cause (if known), and all actions taken — required for HIPAA incident log and CMS reporting if PHI was unavailable for an extended period
Minimum security controls during DDE unavailability:
- Break-glass accounts remain restricted and monitored — do not add them to known-good rules
- Any access using break-glass accounts during the outage is logged per Break-Glass Procedure
- Arctic Wolf MDR monitoring continues — do not suspend alerts during outage
2. Azure PROD Unavailable (Internal Operations Affected)
Internal staff lose access to PROD VMs and tooling. M365 continues independently.
Immediate actions:
- Notify Kevin and Greg — they are the escalation path for domain-level operations
- Verify access via alternative paths: M365 (Teams, email) continues; Keeper accessible
- Determine cause via Azure Portal or Azure Service Health
- Check AWS DR environment — if PROD Azure is down but DDE is up, focus on DDE continuity
Operational constraints during PROD outage:
- Freeze all Terraform changes — do not attempt to apply infrastructure changes during an outage
- Pause scheduled maintenance — rescheduled after restoration
- Continue security monitoring — Arctic Wolf and any agents still running continue; SecOps may be unavailable (it runs in PROD)
- If SecOps is unavailable: agents are configured fail-open — they raise alerts and skip suppression. Expect elevated noise. Do not silence alerts.
Communications during PROD outage:
- Primary: M365 Teams
- If M365 is also affected: mobile phones (see Out-of-Band Communications Plan)
3. Security Monitoring System Unavailable (SecOps Platform Down)
The SecOps platform (ca-secops-prod) may be unavailable independently of broader Azure outages.
Impact: Agents cannot post findings; incident management UI unavailable.
What continues without SecOps:
- Arctic Wolf MDR — fully independent, will page Rory directly for critical incidents
- Azure Monitor alerts — fire independently to IT admin email
- Cortex XDR behavioral detection — continues independently
- Palo Alto threat detection — continues independently
Minimum monitoring during SecOps outage:
- Check Arctic Wolf portal daily for any open incidents
- Check Azure Monitor alert history
- Restore SecOps per SecOps Platform Guide troubleshooting section
- If SecOps will be down for >4 hours, notify Rory to monitor Arctic Wolf and Cortex XDR directly
Minimum Acceptable Security Controls During Emergency Mode
Regardless of what systems are unavailable, these controls must never be suspended:
| Control | Minimum Acceptable State |
|---|---|
| Authentication (MFA) | Entra MFA and Conditional Access remain enforced — do not disable even during outage |
| Break-glass account monitoring | Any break-glass login remains a CRITICAL alert — never suppressed |
| Audit logging | CloudTrail continues in AWS; Azure Monitor logging continues if control plane is up |
| Arctic Wolf MDR | Must not be suspended; coordinate with Arctic Wolf directly if needed |
| Backup immutability | Object Lock and RSV immutability must not be modified during emergency |
| Encrypted storage | No emergency justification to disable at-rest encryption |
If a control absolutely must be suspended for emergency recovery (for example, temporarily adjusting a Conditional Access policy to allow break-glass access):
- Document the change before making it
- Set a specific restoration time — not "when things are stable"
- Restore the control before declaring the emergency resolved
- Log in the risk acceptance register
Communication Plan During Emergency
| Stakeholder | Communication Method | What to Tell Them | Who Tells Them |
|---|---|---|---|
| Medicare customers / patients | Adriana → CMS customer contact | DDE unavailable, estimated restoration, PHI safety status | Adriana |
| CMS (if PHI unavailability > 24h) | Adriana → CMS reporting channel | Outage duration, PHI affected, remediation steps | Adriana |
| Kevin + Greg | M365 Teams → mobile phone fallback | Status update every 2 hours during active emergency | Rory |
| Internal staff | M365 Teams post | What's unavailable, what's working, ETA | Rory |
| HHS (if reportable breach) | Formal HHS breach notification | Per HIPAA Breach Notification Rule (within 60 days for >500 individuals) | Adriana + Rory |
Return to Normal Operations
When the outage is resolved:
Verify systems before re-announcing. Do not declare recovery until you have confirmed DDE patient access works, not just that the Azure status page says green.
Review alerts that fired during the outage. SecOps and Arctic Wolf will have accumulated incidents. Review them — an attacker may have used the outage as cover.
Restore any suspended controls. Any CA policy changes, maintenance modes, or access exceptions granted during the emergency must be reversed.
Document the incident. Log in SecOps or the HIPAA incident log:
- Start and end time
- Cause (if known)
- Systems affected
- PHI availability impact (if any)
- Actions taken
- Follow-up needed
CMS notification review. If DDE was unavailable for an extended period and PHI was inaccessible to patients, consult with Adriana on whether CMS reporting is required.
Emergency Contact Quick Reference
| Person | Role | Contact Method |
|---|---|---|
| Rory | Security Officer, primary decision-maker | Teams → mobile (in Keeper) |
| Kevin | T1 Domain Admin, DR escalation | Teams → mobile (in Keeper) |
| Greg | T1 Domain Admin, DR escalation | Teams → mobile (in Keeper) |
| Adriana | Customer communication, compliance | Teams → mobile (in Keeper) |
| Arctic Wolf | MDR — 24/7 incident support | Arctic Wolf portal → Concierge Security Team; phone in Keeper |
Related Documents
- DR Failover Procedure — technical steps for Azure → AWS failover
- Incident Response — security incident response
- Out-of-Band Communications Plan — when M365 is also unavailable
- Break-Glass Procedure — emergency account activation
- Backup Architecture — RPO/RTO and backup coverage
Document History
| Date | Change | Author |
|---|---|---|
| May 2026 | Initial draft — satisfies HIPAA §164.308(a)(7)(ii)(C). Covers DDE outage, PROD outage, SecOps outage, minimum security controls, communication plan, return to normal. | Rory |