Why procurement measurement goes wrong
Procurement is unusually easy to measure badly. The activity is countable, the value is contested, and the function reports numbers that nobody outside it can independently check. That combination produces two predictable results: a scorecard full of activity, and a benefit figure the rest of the organisation quietly discounts.
Activity metrics are the more damaging, because they change behaviour. Count tenders run and you get tenders that did not need running. Count suppliers rationalised and you get consolidation applied to categories where competition was the only thing keeping the price honest. Count purchase orders processed and you have built a case for keeping a process that should have been automated.
A measure that would look the same whether or not the work was any good is not measuring the work.
The five measures
Five outcome measures cover most of what an executive team needs, and each is defensible because somebody outside procurement can verify it.
- Validated benefitSavings and cost avoidance as separate lines, against definitions finance has agreed, with the baseline stated for each and a recurrent or non-recurrent flag on every line. One number combining the two is the fastest way to lose the credibility of both.
- Contract coverage and controlThe share of addressable spend sitting under a current contract with a named internal owner and a known expiry date. It exposes more risk than any other single measure. It is also argued with, on all three of its terms: what counts as addressable, what counts as current when an expired agreement is still being invoiced against, and whether an owner who has never opened the contract counts. Survive the argument about the denominator and the measure changes the conversation.
- Sourcing cycle timeWorking days from approved requirement to contract award, reported at the median rather than the mean, split between routine and complex. This is the number the business feels, and a long tail is usually why people buy around the process. The mean hides the tail, which is the part that does the reputational damage.
- Supplier performance against contractFor the suppliers that matter, performance against the service levels actually in the agreement rather than against a generic template. Where the contract contains no measurable levels, record that as a known gap rather than inventing one.
- Route complianceThe share of spend that followed the intended buying route. Assume a design problem before a discipline problem: low compliance usually means the compliant route is slower than the alternative. Check both, because some of it is a budget holder with a supplier relationship they would rather not have tested.
Coverage tells you a contract exists. It does not tell you the contract is being applied, and the two are separated more often than they should be. Whether the price charged is the price agreed is checkable, and it is not visible in anything reported at summary level: it needs invoice lines compared against the contracted rate, which is why price non-compliance is usually found during an audit or a system migration rather than by the people responsible for the category. Where the data allows it, run that comparison on the largest agreements as part of the coverage measure rather than as an occasional exercise. It tends to pay for the effort the first time it is run.
Two more belong on the scorecard only where they are genuinely live: social value or sustainability commitments delivered, measured against what the supplier actually undertook at award rather than against ambition, and payment performance where the organisation reports on it or where supplier cash flow is a real risk in the supply base. Supply chain concentration and single points of failure increasingly come up at board level too, and those are supplier questions rather than procurement performance questions.
Agree the definitions, then do the arithmetic
Almost every argument about procurement performance is an argument about definitions nobody settled. Three of them do most of the damage.
- Baseline
- What the comparison is against: previous contracted price at current volumes, the average of the last twelve months, a validated market benchmark, or budget. Different baselines produce different numbers from the same event. The one that gets disputed most is a first-time purchase, where there is no prior price at all.
- Period
- Annual benefit, contract term value, or annualised. A five year deal claimed in full in year one is a forecast wearing a saving's clothes.
- Two flags on every line
- Cash-releasing or not, and recurrent or not. These are the first two things finance looks for and the two most often missing, and in the public sector the first of them decides whether a benefit counts at all.
The reason to be precise is that a benefit number is an arithmetic claim and it is usually examined as one. Take an illustration. A contract runs at a hundred thousand pounds a year. The supplier asks for a five per cent uplift; the settlement is two per cent. Separately, a specification change takes four thousand pounds out of the annual cost. Volumes are forecast to rise three per cent.
- The cost avoided is three per cent of the base, because that is the difference between what was asked for and what was agreed. It is real and it releases nothing.
- The saving is the four thousand pounds from the specification change, and only if the requirement is genuinely still met.
- The budget still goes up, because the two per cent uplift and the volume increase both have to be funded. That is a cost pressure and it belongs on the same page.
- If the four thousand pounds cannot be taken out of the budget line, it is a saving that is not cash-releasing, and finance will treat it accordingly.
Presented as one number, that becomes a seven thousand pound saving, and the first person who unpicks it takes the credibility of the whole scorecard with them. The distinction between the two kinds of benefit is worth understanding properly before setting any target against either.
The measures you do not get to choose
For contracting authorities under the Procurement Act 2023, part of the supplier performance measure is now a statutory floor rather than a design choice. Contracts with an estimated value above five million pounds must have at least three key performance indicators set before the contract is entered into, performance must be assessed against them at least once every twelve months and on termination, and that assessment is published, including a rating where performance is poor. Call-off contracts under a framework are in scope on the same value test.
The practical consequence is not the reporting burden. It is that a measure which will be published has to be one the organisation is prepared to defend, and that a function whose contract management has been nominal will discover that in public.
This is the general position and it is not the whole regime. Scotland has its own rules. A dedicated Answer on the transparency duties is planned rather than spread through the commercial articles, because procedural detail dates quickly and is better maintained in one place.
What breaks a scorecard
- Averages hiding the problemA mean cycle time of forty days can conceal a category where nothing has been awarded in nine months. That category is where people have already stopped using the process.
- Benchmarking against a different organisation's shapeProcurement headcount as a share of spend is not meaningless, but it is unreliable across organisations with different scopes, different levels of outsourcing and different views of what counts as addressable. Your finance director will use it anyway, so normalise it rather than refusing the comparison: state the denominator, exclude transactional heads, and give the contract count alongside.
- A target set before the baseline existsA savings percentage announced before anyone has established what is being spent produces a pipeline built backwards from the target. It misses, and the miss is attributed to delivery rather than to the target setting.
- Definitions agreed once and never defendedA new finance director, a new ledger or a budget gap reopens the question. Get the definitions signed, keep the paper, and expect to have the conversation again at every leadership change.
If the current scorecard is not trusted, the recovery is mostly conversation rather than analysis. Take the three definitions to finance and agree them in writing before producing another number. Establish contract coverage, because it is the measure most likely to be missing and the one that most reliably changes an executive team's mind. Then cut the scorecard to five and retire the activity metrics rather than reporting them below the line, because anything still reported still counts.