Segment before you measure
The most common failure is applying one process to every supplier. A supply base of any size cannot be managed uniformly, and attempting it produces a scorecard exercise that consumes the available time and changes nothing.
Segmentation is a judgement on two axes, and the framing is Kraljic's rather than anyone's original thinking: what the organisation would lose if the supplier failed, and how much value there is in the relationship working better than the contract requires.
- Strategic
- High dependency, hard to replace, or genuinely able to contribute something beyond delivery. Named relationship owner, structured reviews, executive contact on both sides, joint improvement activity that someone is accountable for.
- Important
- Material spend or material operational risk, replaceable with notice. Periodic review against contracted measures, a named contract owner, a clear escalation route.
- Transactional
- The majority by count. Managed by exception, using transaction data that already exists, with someone looking at it when something goes wrong.
Segmentation is not a spend ranking. A low-spend supplier holding a single point of failure belongs above a high-spend supplier in a competitive market. Getting that wrong is how an organisation ends up running quarterly reviews with its stationery supplier and none with the one whose failure would stop a service.
The right size for the strategic list is whatever the team can genuinely run. Each supplier in it means a data pack, a meeting, an action log and the chasing in between, every quarter. A list of a hundred strategic suppliers is a list nobody is managing.
What a scorecard line has to contain
A generic scorecard measures what scorecards usually measure. A useful one measures what was agreed, which is a shorter list and a harder one to argue with. Each line needs the same eight things.
- The measure, and the clause or schedule it comes from.
- The definition, including what is excluded. Availability measured at the service is a different number from availability measured at the component, and the contract will say which.
- The data source, and who produces it. Supplier self-reported, buyer system derived, or jointly reconciled: this single choice determines whether the scorecard is believed.
- The measurement period.
- The target, and the thresholds for a minor and a major failure.
- The consequence, and any cap on it.
- The relief events: what stops the clock, who declares one, and what evidence is required. Most service credit disputes happen here rather than over the measure itself.
- How buyer-side dependencies are handled. Where the supplier's performance relies on your ticket quality, site access or sign-off, those have to be excluded or the score is measuring your organisation.
Measure whatever you like. Remedy only what the contract defines.
That distinction is worth holding firmly, because getting it wrong in either direction is expensive. You can measure anything and share it, and doing so is frequently how a service level gets agreed at the next renewal. What you cannot do is apply a service credit, serve a default notice or withhold payment for a failure against a measure the supplier never agreed to. Presenting unagreed measures as contractual failures is what destroys a scorecard's standing, and withholding payment outside a contractual right is itself a breach.
Where a contract contains no measurable service levels at all, which is common in call-offs and short-form purchase order terms, record it as a known gap and fix it at renewal. In the meantime, delivery timing and invoice accuracy come from your own transaction data and do not depend on the supplier agreeing to anything.
Reviews that change something
A review where the supplier presents their performance and everyone agrees it has been a difficult quarter is a meeting, not a governance mechanism. Four things make the difference.
- The right peopleSomeone from the supplier who can commit to a change, and someone from your organisation who owns the service. Two account managers exchanging reports cannot change anything and both know it.
- Data agreed in advanceThe meeting is for deciding what to do about the performance, not establishing what it was. If the first half is spent disputing the numbers, the data source was never agreed and that is the thing to fix.
- Actions with owners and dates on both sidesIncluding actions on your organisation, which usually exist: late payment, unclear demand, unmanaged internal change. A review that only ever produces supplier actions is not being taken seriously by either side.
- Last period's actions reviewed firstThe cheapest discipline available and the one that most changes how the meeting is treated.
Frequency should follow segment and should separate the two conversations. Operational review of a critical outsourced service is monthly; the commercial review, where term, price and structural change are discussed, is quarterly or twice yearly. Running one meeting for both means the operational detail crowds out the commercial conversation every time.
Whatever the cadence, the performance record has to reach the next sourcing decision rather than being rediscovered during it. A re-tender run without the operational history evaluates the incumbent on their bid rather than on their record, and in a regulated procurement using past performance is possible but constrained: it has to be relevant, evidenced, and handled through the proper route rather than dropped into the evaluation.
Escalation, in the order it has to happen
Performance management works because there is a credible consequence. Where the escalation path has never been used, suppliers act on that information and so do internal stakeholders. The sequence matters, and each step has a contractual character that is easy to skip.
- Log the failure against the clause, in timeIn writing, referencing the specific provision, with the evidence, and within whatever notification period the contract sets. Miss that window and the remedy is often lost regardless of how bad the failure was.
- Escalate operationally under the governance scheduleNamed roles, defined response times. Most contracts have this and most organisations use email instead.
- Serve the right noticeA rectification notice, a default notice and a notice invoking a remediation process are different instruments with different consequences. Serve them as the contract requires, on the person and address it names.
- Run a remediation plan with a defined testA plan period, a specific test, and an express statement that accepting the plan waives neither accrued rights nor the right to terminate.
- Reserve rights, consistentlyIn every piece of correspondence, and avoid a pattern of conduct that looks like acceptance. A year of politely tolerated failures is difficult to convert into a remedy later.
- Apply the consequence, or use persistent breachService credits, volume moved, or termination. Most agreements define persistent breach separately from material breach, as a number of failures within a period, and that is usually the clause that actually gets used, because a single failure rarely reaches the material threshold.
- Plan the exit before you need itTermination is the start of the expensive part rather than the end of the problem. Exit plan, transition assistance, data and asset return, and knowing what the alternative supply looks like before the notice goes out.
None of that works if the people closest to the failure do not know the mechanism exists. The notification window is missed by the service owner far more often than by the contract manager, because the service owner is the one who sees the failure and has never had a reason to read the clause. A walkthrough with the delivery team, built around the situations they actually encounter rather than around the contract's own structure, recovers more value than redrafting anything.
Escalation has a cost and it should not be the first response. But a path that exists only on paper is not leverage, and its absence is usually why a relationship deteriorates for a long time before anyone acts.
Where this is no longer optional
For contracting authorities, performance management on larger contracts has moved from good practice to a published duty, with indicators set before signature and assessments published during the contract's life.
The commercial consequence is larger than the reporting burden, and it changes what a review meeting is. A performance rating that will be published has a value attached to it, because it follows the supplier into their next bid. That makes the supplier considerably more interested in the scorecard, which is useful, and it makes a poorly evidenced measure considerably more dangerous, which is not. It also means the measures have to be written at the point the contract is drafted, by someone who knows what the service will actually be judged on, rather than assembled afterwards from what happens to be measurable.