Skip to content

Operating model

Who does what, and who decides at three in the morning

A SOC fails at its boundaries far more often than at its tooling. This page sets out the three ways the function is commonly organised, the separations that must hold inside whichever you choose, and — the part most often left undecided — exactly who is permitted to act when there is no time to ask.

Models
A · B · C
Hard rule
Separate admin from monitoring
Missing piece
Decision rights

Structure

Three models, and what each costs you

Organisations usually start at A and move to B or C as volume grows. None of the three is wrong — but each has a failure mode, and knowing yours in advance is the difference between managing it and being surprised by it.

A

One team — operations, engineering and response together

A single unit covers monitoring, investigation, detection engineering and incident response, with the roles separated inside it.

Fits whenSmall to medium volume, limited resources, and where speed and compactness matter more than specialisation.

Advantages

  • Fewer handovers, so faster response
  • One shared picture of the situation
  • Cheaper to run and simpler to govern

Risks

  • Burnout — the same people carry every mode of work
  • Conflict of interest between administering the tools and monitoring with them
  • Hard to sustain genuine round-the-clock cover or deep specialist skill
B

SOC separated from incident response

The SOC detects, triages and escalates. A separate response team coordinates containment, recovery and forensics. Detection engineering sits apart from both, or with response.

Fits whenRound-the-clock requirements, several business units or client organisations, high risk, heavy regulation, or high incident volume.

Advantages

  • Clearer accountability at each boundary
  • Better methodological quality and stronger separation of duties
  • Scales without the whole function degrading

Risks

  • Needs more people
  • Introduces handover points, which demand good process and tooling to survive
C

Co-managed — external provider with an internal service owner

Part of the monitoring and investigation, and the out-of-hours cover, is delivered by a provider. What stays inside is the service owner, the integrations, the organisational context, the decisions and the communication.

Fits whenRound-the-clock cover needed quickly with no internal capacity, or a deliberate push to raise maturity over six to twelve months.

Advantages

  • Fast start on continuous cover
  • Access to expertise that would take years to build
  • Often more efficient on cost at low volume

Risks

  • Dependence on the provider
  • Requires contracts and processes that define authority precisely — who may do what, to which systems
  • Accountability stays with you regardless of who does the work

The round-the-clock arithmetic

A single continuously staffed seat needs several full-time people once shifts, leave, training and sickness are accounted for. If the organisation does not have that volume, the honest options are co-managed cover with internal senior tiers during business hours, or an explicit decision that out-of-hours means on-call rather than staffed.

What is not an option is promising continuous cover in a service description and staffing it eight hours a day. That gap is discovered during the incident that needed it.

One catalogue, one owner

Whichever model you pick, keep a single service catalogue and a single accountable service owner. Grey zones between two catalogues are where incidents go to be nobody in particular’s problem.

Separation of duties

What may be shared, and what needs a control point

These separations survive even in a team of four. Where the people must be the same, the roles, the processes and the reviews still have to be distinct — that is what makes it a control rather than a wish.

Continuous monitoring and triage

May be shared

May be a shared function. Requires explicit response-time commitments per severity.

Initial investigation and evidence gathering

May be shared

May be shared, provided there is one case management system and standard playbooks.

Threat hunting

Control point required

Usually sits with the senior tier. Hunting must not crowd out continuous operations — give it its own protected time and its own measures.

Detection engineering and tuning

Control point required

Separate from the duty roster. Rule changes only through change control with peer review.

Tool administration

Control point required

Must not be performed at the same time as an investigation by the same person. Where the people are the same, the roles must be distinct.

Containment actions on live systems

Control point required

Requires pre-authorisation and break-glass procedures. Who may isolate a host, disable an account or change a firewall must be named in advance.

Incident command and communication

Control point required

Communication with management, legal, the data protection function and external parties is centralised through the incident commander.

Forensics and chain of custody

Control point required

Frequently a separate competence, so that methodological integrity and legal value are preserved.

Lessons learned and preventive change

May be shared

May be shared, but the resulting decisions go through the normal change and risk governance.

Compliance reporting and notification

Control point required

The SOC supplies the technical facts. The decision to notify a regulator or a client is taken with governance, legal and compliance.

Decision rights

The question the mandate does not answer

A mandate grants powers to 'the SOC' as a unit. It does not say which individual, at which tier, may isolate a production host at 03:00 without waking anyone. Decompose it, or that decision will be made by whoever feels bravest.

Containment actions

Executed only under pre-agreed playbooks with an explicit approval mechanism. Automatic for the top severities where the conditions are pre-defined; otherwise incident commander approval.

Detection rule changes

Through the change process: a ticket, peer review, a test environment, and a deployment window. Never live on the console during an incident.

Incident communication

Owned by the incident commander, with legal and the data protection function as required. Analysts communicate only through technical channels and defined forms.

Break-glass

A documented emergency path that exceeds normal authority, with mandatory immediate notification and an after-the-fact review. Untested break-glass is not a control.

A worked RACI

ActivityService ownerL1L2Detection eng.Incident responseIT operationsLegal / DPO
Alert triage and validationARCCIII
Incident classificationACRCRCI
Containment — e.g. host isolationAICCRRC
Forensics and evidence collectionAICCRCC
Detection rule changesAICRCII
Threat intelligence feed maintenanceAICRCII
Communication with management or clientAIIIRCR
Post-incident review and improvementACCRRCC

R responsible · A accountable · C consulted · I informed. Exactly one A per row — if you find two, you have found the argument you will have during the incident.

Escalation

Severity, first response, and who gets the decision

Note that the escalation column carries the decision right, not just the recipient. 'Escalate to incident response' without saying who may authorise containment leaves the most time-critical decision undefined.

S1 — critical

15 minutes, round the clock

Active intrusion, ransomware indicators, or a live risk of data exfiltration

Immediately to incident response and IT operations. Containment only with the incident commander, or automatically where pre-agreed conditions are met.

S2 — high

60 minutes

A confirmed incident with limited impact

Incident response within 2 hours. Containment per playbook.

S3 — medium

4 hours in extended-hours cover; 2 hours in continuous cover

A suspicious event that needs more data before it can be judged

L2. Incident response only if confirmed.

S4 — low

1 working day

Noise, informational, or a policy violation

Close, or add to the backlog for detection tuning.

Reconciliation

Four severity scales, one incident

Most SOCs run several severity vocabularies simultaneously — an operational scale, a contractual one, one in the report, and the statutory impact categories — and never map them to each other. That is how an incident ends up rated low internally and notifiable externally.

OperationalService commitmentIn the reportStatutoryWhat the SOC Officer does
S1 — criticalA — criticalHighLikely large / significant — assess immediatelyStart the statutory clock. Notify leadership. Containment under the mandate.
S2 — highB — highHigh or mediumAssess against the criteria — may qualifyRun the significance test explicitly and record the outcome either way.
S3 — mediumC — normalMediumUsually below the threshold — but national rules may still require a reportCheck the national non-large reporting obligation before closing.
S4 — lowC — normalLowInsignificantClose, tune, and include in the periodic aggregate count where required.
Publish this mapping. It is the single artefact that stops the operational scale being mistaken for the regulatory one — and it belongs in the triage procedure, not only in a report.

Rhythm

The cadence that keeps the model honest

A structure on a slide is not an operating model. What makes it real is the set of recurring activities that happen whether or not anything is on fire, each with a named owner.

Per shift

L1 / shift lead
  • Handover: open incidents, watch items, environment changes in flight
  • Confirm detection and log-source health dashboards are green
  • Work the alert queue to the agreed triage time target
  • Escalate against written triggers, not against workload
  • Record the handover before leaving — no verbal-only transfers

Daily

SOC Officer
  • Review incidents opened, escalated and closed in the last 24 hours
  • Check ageing: anything past its severity-linked response target
  • Confirm no log source has silently stopped reporting
  • Clear the escalation decisions that need your authority

Weekly

SOC Officer
  • Tuning review: top noisy rules, and what will be done about each
  • Detection backlog grooming against current threat intelligence
  • Open recommendations to system owners — chase what has slipped
  • Rota and coverage check for the coming two weeks

Monthly

SOC Officer
  • Issue the management report: posture, incidents, KPIs, risks
  • Review MTTD and MTTR trends and explain any movement
  • Close out post-incident review actions that are now due
  • Review coverage changes: new systems onboarded, sources retired

Quarterly

SOC Officer + CISO
  • Exercise: table-top or technical, with a written result
  • Efficacy test — does the detection stack still catch what it claims?
  • Review the SOC's own risks: key-person, capacity, tooling contract
  • Refresh the threat model and re-prioritise the detection roadmap

Annually

Management body
  • Re-approve the SOC mandate, scope and delegated authority
  • Reassess the effectiveness of the risk-management measures
  • Confirm management-body cybersecurity training is current
  • Review staffing, competency matrix and succession for each tier

Risks

Close these on day one

Each of these is predictable, and each has a known control. Deciding them early costs a meeting; discovering them late costs an incident.

RiskControls
Unclear accountability — 'who was supposed to do that?'A RACI matrix, a published service catalogue, a named incident commander role, and written escalation rules.
Conflict of interest — the same person writes the rules and judges their qualityPeer review and change management; a separate detection engineering role.
Burnout and an always-on cultureA published shift plan, rotation, realistic response commitments, on-call compensation, and automation of the repetitive work.
Provider risk — privileged access and data confidentialityContractually defined rights and limits, audit rights, confidentiality obligations, least privilege, and break-glass procedures.
Data fragmentation across separate tools and dashboardsOne case management system, standardised enrichment, and centralised log onboarding.
Alert fatigueUse-case prioritisation, systematic tuning, intelligence quality filtering, and measuring false positives per rule.

Reference

What to build on

You do not need to invent an operating model. These are the references the profession actually uses — cite them when a design decision is challenged, and you are arguing from established practice rather than from preference.

  • ENISA — How to set up a CSIRT and SOC

    The European reference for standing up the function, and for the distinction between a SOC and a response team.

  • NIST SP 800-61 — Computer security incident handling

    The incident lifecycle, and the basis most response procedures are derived from.

  • NIST SP 800-137 — Information security continuous monitoring

    Designing what to monitor and how continuously, rather than monitoring what is easy.

  • ISO/IEC 27035 — Information security incident management

    The management-system view of incident handling, useful where a certification is in scope.

  • NIST Cybersecurity Framework

    The common language for explaining to a board where the SOC sits among the wider security functions.