When Microsoft 365 fails, most of the damage comes from confusion and delay rather than the outage itself. Three scenarios account for nearly every M365 emergency a small business faces, and each needs a different response. This playbook covers all three.
If your team can't send email, open shared files or reach clients, every minute costs revenue and trust. The goal in the first half hour is operational continuity, not perfect diagnosis.
Don't assume a full outage. Check whether it affects one user, one location or everyone; test Outlook Web App and mobile access; check Microsoft 365 Service Health if you can reach it; and verify your own internet, DNS and firewall. If every user and every access method fails, treat it as P1.
Name an incident lead. Open one internal channel and keep all traffic there. Pause non-essential IT changes — parallel troubleshooting during an incident creates new faults. Start a timestamped log.
Use one update format throughout: what failed, what's impacted, when the next update comes.
Most businesses badly under-communicate during outages. Send three messages quickly: an internal staff advisory, a customer-facing delay notice if needed, and a leadership snapshot with impact and next update time. Keep them calm and free of technical detail.
Resist broad configuration changes while the cause is uncertain. More incidents are extended by speculative fixes than by slow diagnosis.
Route urgent communication to phone or SMS. Use a backup shared mailbox or channel if one exists. Prioritise the teams handling revenue and customer response. If the outage extends, move to an hourly update cadence.
When escalating to your provider or Microsoft, include exact start time, affected user count, affected services, error messages or screenshots, and what you've already tested. Clean evidence cuts resolution time substantially — vague escalations get queued.
The usual reports: Outlook repeatedly asking for a password, sign-in working in a browser but failing in desktop apps, intermittent MFA loops, or mailbox access working for some users and not others.
The pattern tells you where to look:
Stabilise first, optimise second. Restore access safely, capture the timeline and policy state at the time of the incident, confirm there's no broader regression, then log the root cause. Never permanently weaken a policy under pressure — temporary exceptions have a way of becoming permanent.
Admin lockout can freeze email, Teams, file access and security controls within minutes, and it's the most disruptive of the three.
Confirm whether the lockout is account-specific or tenant-wide. Check whether any secondary admin still has access. Freeze non-essential changes and privilege updates. Capture exact error messages and timestamps. If no admin account is available at all, escalate as a P1 identity incident immediately.
1. Secondary admin. If another admin can sign in: reset the affected credentials, re-register MFA for the locked account, and check whether conditional access is contributing.
2. Break-glass account. If you maintain emergency accounts, use one under a controlled process, restore minimum admin operations only, and rotate the credentials immediately afterwards.
3. Microsoft support. If all admin access is blocked, open a high-priority case with your tenant domain, an impact statement and the incident start time. Keep a single owner for support communication — multiple people chasing the same case slows it down.
Lockouts are almost always avoidable and almost always traceable to the same causes: a single global admin dependency, gaps in the MFA reset process, conditional access conflicts, an admin account tied to one person's personal phone, or no documented emergency access path.
Within 24 hours: verify Exchange, Teams, SharePoint and Intune; audit admin role assignments; review sign-in and risky sign-in logs; rotate any credentials used during response; and document the timeline and decisions while they're fresh.
Most of this is cheap to set up and expensive to skip:
Escalate quickly when impact spreads beyond one team or site, when executives or customer-facing staff are blocked, when a third auth incident occurs within 30 days, or when policy interactions become too complex to isolate safely. Repeated incidents mean the process needs redesign, not just another fix.
If your business runs on Microsoft 365 for client communication and operations, these are continuity risks rather than IT inconveniences. We run outage-readiness and admin resilience reviews, and our 24/7 coverage means overnight incidents get a person rather than a voicemail.
Our free IT health check reviews your Microsoft 365 configuration, admin resilience and recovery readiness, with a written report you keep either way.