A security plan can look convincing in a board pack and still fail at the point of decision.
The same is true of training records, audit scores and completed exercises. If you want to know how to measure security capability, the question is not whether those things exist. It is whether people can recognise a developing problem, make sound decisions and carry out their role when time, information and confidence are limited.
That distinction matters. Organisations often measure activity because it is easy to count. They report courses completed, policies reviewed, incidents logged and equipment installed.
These are useful management controls, but they are not direct proof of capability. They describe what has been done, not what a team can reliably do.
Start with the outcome that matters
Security capability is the practical ability of people, systems and leadership to reduce risk and respond effectively. It is demonstrated through judgement, coordination and action. A capable team does not merely know its procedure. It understands why the procedure exists, when it applies and when circumstances require escalation or adaptation.
Measurement should therefore begin with the operational outcome you need. For a venue, that may be the ability to identify concerning behaviour, report it clearly and make proportionate decisions without disrupting normal operations unnecessarily. For a corporate security team, it may be the ability to assess a credible threat, advise senior leaders and coordinate a response across several sites. For a project manager, it may mean maintaining security controls when programme pressure, contractors and changing site conditions compete for attention.
Broad statements such as ‘improve awareness’ are not measurable. Define the expected performance instead. Who must do what? What information should they recognise? Who do they need to involve? How quickly must a decision be made? What does an acceptable outcome look like?
This is not an argument for reducing every security task to a stopwatch. Judgement cannot be measured by speed alone. A fast but poorly founded decision can create more risk than a slower, properly escalated one.
The point is to set standards that reflect the role and its consequences.
Measure security capability through behaviour
Knowledge checks have a place.
They can establish whether someone understands terminology, reporting routes or core principles. On their own, however, they are a weak measure of operational readiness. A person may select the correct answer in a quiet online assessment and still hesitate, miscommunicate or make an unsafe assumption in a live situation.
The better test is behaviour. Observe how people apply knowledge to realistic, role-relevant situations. This does not require elaborate simulations for every role. A well-designed discussion based on a credible scenario can reveal a great deal: what someone notices first, what they dismiss, how they prioritise, whether they seek clarification and when they escalate.
For example, a front-of-house supervisor may be given a situation involving unusual activity, incomplete information and a busy public environment. The assessor is not looking for a rehearsed speech. They are looking for the quality of the supervisor’s observation, the clarity of their reporting, their understanding of authority and their ability to protect people while avoiding unnecessary disruption.
For senior decision makers, the focus changes. They may need to demonstrate how they weigh competing advice, communicate priorities, maintain oversight and avoid becoming a bottleneck. Leadership capability is often exposed when information is imperfect and the operational picture changes.
That is where paper-based assurance is least reliable.
Use more than one source of evidence
One assessment rarely gives a complete answer. A written test can show understanding. An observation can show applied behaviour. Interviews can expose assumptions and gaps in confidence. Exercise debriefs can show how the team communicates and learns. Incident reviews, where available, reveal what happened when the pressure was real.
Each source has limitations. Observations can be affected by nerves or by an unusually simple scenario. Incident data may be too sparse, or may reflect reporting culture more than actual performance. Self-assessment is valuable for reflection but frequently overstates competence, particularly where people have not been exposed to the demands of the role.
Triangulating evidence reduces those weaknesses. If a team scores well in knowledge evaluation but repeatedly struggles to provide concise reports during exercises, the issue is not simply training completion. It may be communication discipline, role clarity or confidence under pressure.
If leaders describe an effective escalation process but staff cannot explain who makes the final decision, the control exists in theory rather than practice.
Mildot Group’s capability evaluations are built around this principle. Immediate feedback has value because it identifies where a practitioner’s judgement or knowledge needs development, rather than merely confirming that an assessment has been completed.
Build a scoring model people can understand
A scoring model should support better decisions, not create a false impression of precision.
Avoid a single percentage that combines unrelated strengths and weaknesses into a reassuring average. A team may be strong at access control and weak at identifying behavioural indicators. An overall score can hide the weakness that matters most.
Assess capability across a small number of areas linked to the role. These may include threat recognition, risk assessment, reporting, decision making, coordination, leadership and recovery. Score each area against clear descriptions of performance.
A four-level scale is often enough. At the lowest level, the person requires close direction and may not reliably identify the issue. At the next level, they can follow a defined process in routine circumstances. A competent practitioner can apply the process, explain their reasoning and escalate appropriately when conditions change. At the highest level, they can lead others, manage ambiguity and improve the team’s response.
The descriptions matter more than the number. Two assessors should be able to explain why someone was judged competent and point to evidence.
If they cannot, the score is too subjective to guide investment or assurance.
Capability thresholds must also reflect risk. A security manager responsible for a complex site should not be measured against the same standard as a colleague whose duty is simply to identify and report.
Equal scoring systems for unequal responsibilities create misleading assurance.
Test the voids that matter most
A good capability assessment does not attempt to test everything at once.
It starts with the consequences of getting something wrong. Ask where poor judgement, delay, confusion or weak coordination would cause the greatest harm. Then test the decisions that sit at those points.
For organisations preparing for Martyn’s Law, this means looking beyond whether staff have completed awareness training. Can relevant personnel identify concerns and communicate them effectively? Do managers understand their responsibilities? Can teams work together during disruption? Is there confidence in the practical arrangements that sit behind the plan?
Tests should be proportionate and safe. There is no need to create distressing or operationally sensitive scenarios to assess whether reporting routes, authority and decision making work. The purpose is improvement, not theatre.
A realistic assessment creates enough uncertainty to reveal thinking, while remaining controlled and relevant to the workplace.
Do not overlook routine pressures. Security failures are often caused less by dramatic events than by fatigue, competing priorities, poor handovers, contractor churn or the assumption that somebody else has dealt with an issue.
Capability measurement should examine how the organisation performs during normal business conditions, because that is where standards tend to drift.
Turn findings into development, not a filing exercise
Measurement without follow-through is simply another form of paperwork.
Every assessment should lead to a practical decision: maintain the standard, coach an individual, revise a process, clarify accountabilities, improve an exercise programme or obtain specialist support.
The response must match the gap. If staff do not know a reporting route, targeted instruction may be enough. If they know it but fail to use it under pressure, they need practice and feedback. If several teams interpret responsibility differently, the problem is organisational design, not an individual training deficit. Sending everyone on the same course is a common but lazy response.
Set a review point and measure again. Capability changes with staff turnover, restructures, new sites, altered threats and operational pressure. An annual assessment may be appropriate for some functions, but high-risk roles and changing environments need more frequent checks. Reassessment should test whether the improvement has transferred into behaviour, not merely whether attendance was recorded.
The uncomfortable truth is that capability cannot be claimed. It has to be demonstrated, repeatedly, in conditions close enough to reality to make the evidence meaningful.
The most useful question for any security leader is not Are we compliant? but What would our people actually do next?
.
Useful Links:
.
