Six Reasons Your Supplier Scorecard Is Inflating Every Score




Pull up the last supplier scorecard your business completed. Look at the final number.

If it sits somewhere between 70 and 85 percent, and if the previous four suppliers you scored also landed somewhere between 70 and 85 percent, you do not have a measurement system. You have a formatting exercise that produces a comfortable number regardless of what the supplier actually did.

This is not a rare failure. Most supplier scorecard mistakes are arithmetic, not judgement, and they are baked into the spreadsheet before anybody starts rating anything. The people using it are doing their jobs properly. The file is quietly working against them.

Here are the six that turn up most often, roughly in order of how much damage they do.

One: the maximum score is not actually the maximum

This is the big one, and it hides well.

A scorecard has a column for the score you gave and a column for the highest score available. Divide one by the other and you get a percentage. Simple enough — except that in a surprising number of spreadsheets, the “maximum” column is calculated wrong.

GET THE LATEST SUPPLIER EVALUATION SCORECARDS HERE

The version I see most often comes from yes/no questions. A “yes” is worth five points. A “no” is worth zero. But the maximum column has been built with a formula that returns three for a “no” instead of five, usually because somebody copied a formula from a differently-scaled section and never checked it.

The effect is that a supplier who answers no to everything does not score zero. They score around a third. And every supplier above them is lifted by roughly the same margin. Your worst supplier looks mediocre, your mediocre supplier looks acceptable, and nobody ever gets a score bad enough to trigger the conversation the scorecard exists to trigger.

The test takes two minutes. Fill in the worst possible answer on every line and see what the total says. If it is not zero, your denominator is broken.

Two: a flat average treats everything as equally important

Four categories: cost, delivery, quality, responsiveness. Each produces a percentage. The overall score averages the four.

That is a policy statement, and it is almost certainly not your policy. It says a supplier who ships defective parts on time and answers the phone politely is equivalent to one who ships perfect parts and is occasionally slow to reply. In an aerospace or medical supply chain that is not merely wrong, it is the kind of wrong that ends up in a regulatory finding.

Categories need weights, the weights need to live in a visible cell rather than buried in a formula, and somebody senior needs to have agreed them. The useful side effect is that the argument about weighting happens once, in a room, rather than repeatedly and implicitly every time somebody reads a score.

Make the weights sum to one hundred percent and put a check on the cell that turns red when they do not. Somebody will eventually edit one without adjusting the others.

Three: “not applicable” is being counted as zero

You built the scorecard for a machining supplier. You are now using it on a logistics provider. Eleven of the questions are about process capability and tooling control, and none of them apply.

If your spreadsheet treats a blank or an N/A as zero, that logistics provider is being penalised for not owning a CMM. Their score drops, they push back, and the honest answer is that the scorecard is wrong rather than the supplier. So somebody starts leaving questions blank, or marking them generously, and the whole thing stops meaning anything.

The correct behaviour is that an N/A line contributes nothing to the score and nothing to the maximum. It disappears from both sides of the division. The supplier is scored only on what was relevant to them.

This one detail is what makes a single scorecard usable across a mixed supply base, which is worth more than it sounds. The alternative is maintaining nine slightly different versions and discovering, eighteen months later, that four of them have diverged in ways nobody can reconstruct.

Four: the ratings are free text

Somebody built the scoring with a formula that checks whether the rating cell says “Don’t Know/Not Rated”. Somewhere else in the same workbook, another formula checks for “don’t know, not rated” with a comma. There is no dropdown, so the person filling it in types “N/R” or “unknown” or leaves it blank.

None of those match. The formula falls through to its default, which is usually zero, and the score is wrong in a way that produces no error message and no visible symptom.

I have opened commercial supplier scorecards where two categories were only scoring correctly by accident, because a branch that never fired happened to fall through to a value that was coincidentally right. That is not a system anyone should be making sourcing decisions on.

Every rated cell gets a dropdown. Not as a nicety, but because a scoring engine that depends on how somebody spells things is not a scoring engine.

Five: the band legend does not match what the sheet produces

The bottom of the scorecard says: 90 and above is a preferred supplier, 70 to 89 is approved, 55 to 69 is conditional.

The cell above it says 0.758.

Nobody has ever mapped one to the other, because the score is calculated as a fraction and the bands were written as whole numbers, probably by a different person, probably in a different year. So the verdict gets assigned by eye, which means it gets assigned by whoever is in the room and how they feel about the supplier that morning.

Fix the formatting so the score and the thresholds are the same unit. Then have the sheet write the verdict itself, so it is not a judgement call made under mild social pressure with the account manager sitting opposite you.

And once your denominator is honest, expect scores to drop across the board. Recalibrate the bands to match. If you keep your old thresholds after fixing the maths, you will suddenly have a supply base full of failures and a lot of very confused people.

Six: one scorecard is being asked to do five different jobs

The questions you ask before awarding a contract are not the questions you ask at the quarterly review. Before the award you want financial standing, capacity, references, certification scope, tooling ownership. At the review you want on-time delivery, defect rate, corrective action response time, invoice accuracy — things that only exist once the relationship is running.

Most organisations have one scorecard and use it for both. So half the questions are unanswerable in any given context, and everyone gets into the habit of skipping lines, which trains people to treat the whole document as approximate.

Audit questions are different again. Risk assessment is different again. IT and SaaS vendors need questions about data residency, breach notification and exit terms that have no equivalent in a machining supplier’s evaluation. Contractors on your site need safety weighted far above everything else, because that is where the exposure actually is.

Separate documents for separate decisions, and a way to roll the results together when you want the portfolio view.

What to do this week

You do not need to rebuild anything to find out where you stand. Open your current scorecard and run four checks.

Fill in the worst possible answer everywhere and confirm the total reads zero. Check whether your overall score is a weighted calculation or a flat average. Mark a question N/A and confirm the maximum drops by the same amount as the score. Then read the band legend and check it is expressed in the same unit as the cell above it.

If all four pass, your scoring engine is sound and any problems you have are about discipline rather than arithmetic — which is a much better problem to have.

If they do not pass, then every supplier evaluation you have filed for the last few years is somewhere between slightly optimistic and actively misleading, and the fix is a morning’s work rather than a project.

GET THE LATEST SUPPLIER EVALUATION SCORECARDS HERE