Return to Learn

The MSP SLA Checklist Toronto SMBs Should Run Before Signing

The MSP SLA Checklist Toronto SMBs Should Run Before Signing

Most SMBs only discover their SLA is weak during an incident. On paper the provider promised “fast response” and “proactive support”; in practice there are delays, unclear ownership, and exceptions buried in the contract. By month three, the gap is obvious.

When you're comparing IT providers, price is not the biggest risk. A weak Service Level Agreement is. A strong one protects both sides — your team gets predictable outcomes, your provider gets clear priorities and boundaries. This checklist is for Toronto and GTA business owners evaluating managed IT proposals.

Response, mitigation and resolution are three different things

Most providers emphasise first response time because it's the easiest target to hit. But business impact depends on when the issue is actually fixed. A contract guaranteeing only response can leave you waiting days for real recovery while every SLA target is technically met.

Insist all three are defined:

  • First response — acknowledgement and work begins
  • Mitigation — temporary containment, service usable again
  • Resolution — full fix and root cause addressed

A priority model with real definitions

Without written severity definitions, “urgent” becomes subjective — and it will be defined in the provider's favour at 2am. Require explicit examples tied to your environment:

  • P1 Critical — total outage, major security incident, core operations blocked
  • P2 High — partial outage or high-risk degradation, multiple users blocked
  • P3 Medium — single-system issue, workaround available
  • P4 Low — minor defect, request or change with no urgent impact

Every priority needs its own targets and its own service window.

The 12-point checklist

1. Scope of covered services

Confirm what's in: help desk, endpoint support, Microsoft 365, network, vendor coordination, security operations. Then confirm in writing what is not included. The exclusions matter more than the inclusions.

2. Support hours and after-hours rules

State exact hours in Eastern Time, coverage level by priority outside those hours, holiday handling, and whether after-hours is included or billable. If your business runs evenings or weekends, business-hours-only support is operational risk dressed as a saving.

3. First-response targets by priority

A reasonable baseline: P1 within 15–30 minutes, P2 within an hour, P3 within four business hours. Written down, not implied.

4. Mitigation and resolution targets

Separate targets for service restored versus root cause resolved. Without the second, recurring problems never get fixed — they just get closed repeatedly.

5. Escalation matrix, contractual not informal

Require named tiers (L1/L2/L3 plus management), a maximum time before forced escalation on unresolved P1 and P2, and explicit ownership of vendor escalations to Microsoft, your ISP and cloud providers. Without this, high-impact incidents stall in queueing loops while everyone waits for someone else.

6. Communication standards during incidents

Update frequency (every 30–60 minutes for P1 is normal), the channel, and who receives executive updates. How a provider communicates during an outage tells you more about them than any reference call.

7. Patch and maintenance commitments

Patch cadence by asset type, a defined window for critical vulnerabilities, maintenance windows, and exception handling.

8. Security baseline and ownership

Security belongs inside the SLA, not sold as optional project work. Document who owns MFA enforcement, endpoint protection health, privileged access controls, backup alerting and suspicious sign-in monitoring — plus the incident notification timeline and how quickly you receive a post-incident report. Vague security deliverables increase your exposure while reducing their accountability.

9. Backup and recovery obligations

Backup success alone is meaningless; recovery performance is what matters. Require backup frequency by workload tier, RPO and RTO by system, restore test cadence with reporting, and named ownership for failed backups and failed restore tests. A backup policy without tested restores is not a backup policy.

10. Change management and planned maintenance

Approval process for production-impacting changes, rollback requirements, post-change validation, and notice period for planned work. Critical if your internal team is lean — which, if you're hiring an MSP, it is.

11. Reporting and governance cadence

Monthly operational report, quarterly business review, KPI trends (ticket volume, repeat incidents, SLA adherence, security posture), and an action register with owners and due dates. A good provider reduces repeat tickets over time and can show you the graph.

12. Exit and transition terms

Even good partnerships end. Define notice period, documentation handoff, admin credential transfer, and transition support hours before you sign. Ownership of your documentation and credentials should never be ambiguous.

Commercial terms that quietly carry risk

  • “Unlimited support” paired with broad exclusions
  • Ambiguous boundaries between project work and support
  • Mandatory multi-year terms with no performance-based exit
  • Termination language with excessive lock-in
  • Unclear ownership of documentation and credentials

Ask for SLA credits or a remedy model tied to repeated breaches. If a provider resists clear language, treat that as the answer.

A weighted scorecard for comparing providers

Comparing proposals on monthly cost alone is how businesses end up switching again in eighteen months. Score them instead:

  • Incident response and resolution model — 25%
  • Security controls and accountability — 20%
  • Backup and recovery maturity — 20%
  • Governance and reporting quality — 15%
  • Commercial fairness and transparency — 10%
  • Local responsiveness and communication — 10%

The on-site question, specific to the GTA

Remote support handles most issues, but not all. Clarify on-site response expectations for critical incidents, travel coverage boundaries across GTA regions, and who owns hardware and vendor dispatch. Local presence improves outcomes only when the expectations are documented.

Five common mistakes

  1. Choosing the lowest monthly cost without reading the exclusions
  2. Accepting generic SLA language with no priority-specific targets
  3. No defined owner for escalations
  4. No security ownership matrix
  5. No written restore expectations

Five-minute self-audit of your current agreement

If you can't answer these from your existing contract, your SLA needs work:

  • Do we have response and resolution targets by priority?
  • Is after-hours support explicitly documented, and is it billable?
  • Do we know which security controls are included and who owns them?
  • Are backup restores tested on a defined schedule, with results reported?
  • Is there a written escalation path with an update cadence?

If your provider closes tickets quickly but the same issues keep returning, the SLA is operationally weak regardless of what the metrics say.

The takeaway

An MSP contract should be more than a support promise — it's an operating agreement with accountability attached. The best SLA is specific, measurable and enforceable, and the effort spent getting the language right is far cheaper than remediation during an outage.

MapleOps will review your current SLA and flag gaps in response targets, security coverage and recovery readiness. Our free IT health check includes an SLA review — you get a written report either way.

Related reading