A major IT incident is one of the clearest tests of an organisation’s technology operations. When a business-critical system fails, employees need to know how to continue working, customers expect reliable updates and technical teams must coordinate quickly across internal and external providers.
In these moments, a service level agreement is more than a document stored in a supplier portal. It provides an agreed framework for response, communication and restoration. Clear service levels help organisations act with purpose when time, information and attention are limited.
Table of Contents
Set Expectations Before an Outage Occurs
During an outage, uncertainty can be as damaging as the disruption itself. If no one knows who owns the next action, how quickly users should receive an update or when a supplier must become involved, valuable time can be lost.
A well-designed SLA establishes these expectations in advance. It should define the services covered, the support hours available, the severity levels used to classify incidents and the response and resolution targets that apply.
For example, a failure affecting an online payment process may require immediate escalation and frequent updates, particularly during peak trading hours. A fault with a non-essential internal reporting tool may still need investigation, but it is unlikely to require the same urgency or level of communication.
Define What Counts as a Major Incident
Teams cannot respond consistently if “critical” means something different to every stakeholder. A clear priority model makes it easier to assess impact and trigger the right process.
Consider Impact and Urgency Together
A major incident is usually defined by its impact on important business activity. Useful considerations include:
- The number of users, customers or locations affected
- Whether revenue-generating or regulated activity has stopped
- Whether there is a practical workaround
- The risk to security, data or compliance
- The time sensitivity of the affected service
- The potential reputational impact
A payroll application that fails shortly before a processing deadline may be critical even if only a small finance team uses it. Conversely, a problem affecting a large number of users may be less urgent if an alternative service is available.
Make Communication Part of the Service Commitment
Technical restoration is essential, but people also need clear, accurate information while the issue is being resolved. Poor communication can lead to duplicate support requests, conflicting messages and frustration among users who do not know whether anyone is addressing the problem.
Agree Update Timelines
Service levels should specify how and when updates are shared during a serious incident. This might include:
- An initial acknowledgement within a defined period
- A named incident coordinator or service owner
- Regular progress updates for affected users and stakeholders
- A clear statement of known impact and available workarounds
- Notification when service is restored
- A follow-up review for significant disruptions
Updates do not need to include every technical detail. They should explain what users need to know: which service is affected, what action is underway and when the next update will be provided.
Keep Messages Consistent
A single communication owner can help prevent mixed messages. The service desk, technical teams, suppliers and senior stakeholders should work from the same confirmed information, especially when the cause of the incident is still being investigated.
Clear communication also protects support teams from being overwhelmed. If employees can see that an issue is already known and being handled, they are less likely to raise separate tickets for the same problem.
Clarify Responsibilities Across Suppliers
Many modern services depend on several parties. An organisation may use a cloud host, software provider, network partner, identity platform and internal IT team to deliver one user-facing service.
Without agreed responsibilities, each party may investigate only its own component while nobody coordinates the wider service recovery.
An effective SLA should make clear:
- Who owns the end-to-end service experience
- Who declares a major incident
- Which team leads technical coordination
- How external suppliers are engaged and escalated
- Who communicates with business stakeholders
- When the incident can be considered resolved
This shared understanding is particularly valuable when the technical fault sits outside the organisation’s direct control. Even if a supplier is working on the underlying issue, the business still needs one accountable team to manage communication and recovery.
Review Performance After Restoration
Restoring service is the immediate priority, but the work should not end there. A post-incident review gives teams an opportunity to understand what happened, whether service targets were met and what could prevent a recurrence.
The review should examine the timeline, user impact, communication quality, escalation process and root cause. It should lead to specific improvement actions, such as better monitoring, updated support procedures, stronger supplier escalation routes or changes to capacity.
The purpose is to learn, not to assign blame. If a target was missed because the agreed support model no longer reflects the service’s importance, the SLA itself may need to be updated.
FAQs
What is the role of an SLA during a major incident?
An SLA defines expected response times, responsibilities, communication arrangements and restoration targets. It helps teams coordinate their actions and gives users clearer expectations during disruption.
Should every service have the same incident targets?
No. Targets should reflect business impact. Customer-facing, revenue-critical or regulated services often need faster response and restoration commitments than lower-impact systems.
What should an incident update include?
It should explain the affected service, the known user impact, any available workaround, the action being taken and the timing of the next update.
What happens if a service-level target is missed?
The organisation and provider should review the cause, assess the business impact and agree improvement actions. Repeated misses may indicate a need to change the process, capacity or service agreement.
Conclusion
Service levels are most valuable when a serious disruption occurs. By defining priorities, communication expectations and shared responsibilities before problems arise, organisations can respond more calmly and consistently. A practical SLA supports faster recovery, clearer accountability and a more trustworthy experience for the people who depend on essential IT services.

