What the Vendor Evaluation Scorecard does
This scorecard turns weighted criteria into a ranking, and shows every step of the arithmetic: each criterion's share of the total weight, each score normalised onto 0-1 by the scale you set, and each contribution to the final figure. There is no model, no adjustment and no hidden number - if you disagree with the result, change a weight or a score and watch the total follow.
It also does the thing most scoring templates leave out: it tells you how fragile the answer is. For every criterion it computes the weight at which the leader and the runner-up swap places. When a two-point shift in one weight reverses the decision, the decision is not really about the score, and that is worth knowing before the award letter goes out.
How to use it
- List the criteria as
criterion, weight, direction, minimum. Direction isbenefitwhen a higher score is better orcostwhen a lower one is; the minimum is an optional gate. Weights need not add to 100 - they are normalised. - Set the scale you are scoring on. A 1-10 scale and a 0-5 scale give identical rankings, because scores are normalised, but keep one scale for the whole exercise.
- Enter one vendor per line with a score for every criterion, in the same order.
- Read the ranking, then the contribution table: it shows both the raw score and the points it contributed.
- Check the sensitivity table before deciding. If the top two are close and one weight decides it, record why that weight is right.
Reading the results
The total is out of 100 and is the sum of weight share x normalised score. A vendor scoring the maximum everywhere gets 100; the minimum everywhere gets 0.
A contribution shows what a criterion actually did for a vendor. A criterion with 10% of the weight cannot contribute more than 10 points, however good the score.
A failed gate is reported but does not change the arithmetic. Whether a minimum is genuinely mandatory is a procurement decision, not a calculation.
"Weight that flips it" is exact, not a search: the difference between two vendors is linear in any single weight, so the crossing point is solved directly. "Never" means no weight for that criterion alone reverses the top two.
Worked example: four criteria, two close vendors
Price 40, quality 30, delivery 20 and support 10, scored 1-10. Vendor A scores 8, 6, 7, 9 and Vendor B scores 6, 9, 8, 5.
Normalising each score as (score - 1) / 9 and weighting gives Vendor A (0.4x7 + 0.3x5 + 0.2x6 + 0.1x8) / 9 = 6.3/9 = 0.700, and Vendor B (0.4x5 + 0.3x8 + 0.2x7 + 0.1x4) / 9 = 6.2/9 = 0.689. A leads by 1.1 points out of 100.
The sensitivity table shows why that is not a decision. Reduce price's share of the weight from 40% to 36.8% - about three points - and the two are exactly level; below that, B leads. The whole result rests on price being worth four tenths of the evaluation rather than a third.
Change price to a cost criterion instead, so a lower score is better, and the answer reverses outright: A falls to 4.3/9 and B rises to 5.8/9. The direction of a criterion is not a detail.
Formulas and scoring rules
- Weight share
share = weight / sum of all weightsSo weights of 4, 3, 2, 1 behave exactly like 40, 30, 20, 10.- Normalised score, benefit criterion
n = (score - scale minimum) / (scale maximum - scale minimum)- Normalised score, cost criterion
n = (scale maximum - score) / (scale maximum - scale minimum)- Weighted total
total = sum over criteria of share x nShown as a percentage out of 100.- Sensitivity crossing point
solve w x (dk - dOther/otherShare) + dOther/otherShare = 0 for wdk is the difference in normalised scores on that criterion; dOther is the weighted difference across the rest. Linear, so the crossing point is exact.
Why weights should be set before the scores
A scorecard is only a decision aid if the weights were agreed before anybody saw how the vendors performed. Set them afterwards, and the exercise becomes an elaborate way of justifying a choice already made - and the sensitivity table will show exactly how small the adjustment had to be.
The practical discipline is to publish the weights with the tender, score against them, and then run the sensitivity check. If the result is robust across a reasonable range of weights, the score is doing real work. If it is not, say so in the decision note rather than hiding behind a decimal point.
Gates, and the difference between a threshold and a weight
A weight says "this matters more". A gate says "below this, nothing else matters". They are different instruments and mixing them up is a common tender error: giving safety accreditation a high weight lets a vendor trade it away against price, which is almost never the intention.
So gates here are reported separately and never alter the arithmetic. A vendor that fails a mandatory minimum is excluded by a person, with a reason, not quietly penalised by a formula.
Limitations: what the result does not prove
- The arithmetic is exact; the inputs are opinions. Two evaluators scoring the same submission will differ by a point or two, which is often larger than the gap between vendors.
- It cannot check whether the criteria are the right ones, whether they overlap, or whether one of them is a proxy for another. Overlapping criteria double-count.
- It is not a procurement process. Regulated tenders have rules about publishing criteria, moderation, record-keeping and challenge periods that no calculator can satisfy.
- It scores capability; the Vendor Quotation Comparator prices cost. Neither is complete on its own.
Privacy: where your data goes
Everything you paste, type or drop is processed in this browser tab. It is not uploaded, logged, stored or sent to analytics. Session recording and tag-manager scripts are switched off on this page.
Standards and sources
- ISO 20400 - Sustainable procurement guidance - checked 19 Sep 2026
- ISO 9001:2015 - Quality management systems (control of external providers)
Frequently asked questions
How do I build a weighted vendor scorecard?
Agree the criteria and their weights before scoring, choose a single scale, score each vendor on each criterion, then multiply each normalised score by its weight share and add them up. This page does the arithmetic and shows every contribution, so the total can be reconstructed by hand.
Do the weights have to add up to 100?
No. They are normalised by their own total, so 4, 3, 2, 1 gives exactly the same result as 40, 30, 20, 10. Using round percentages is a presentation choice, not a requirement.
What is the difference between a benefit and a cost criterion?
For a benefit criterion a higher score is better, so it is normalised as (score - minimum) / range. For a cost criterion - a day rate, a lead time - a lower score is better and the normalisation is inverted. Getting the direction wrong reverses the ranking outright.
What does the sensitivity table tell me?
The weight at which the leader and the runner-up change places, keeping the other weights in proportion. It is solved exactly rather than searched for. If a three-point shift in one weight flips the result, the score is not what is deciding the award.
Should a mandatory requirement be a criterion with a high weight?
No. A high weight still lets a vendor trade the requirement away against price. Use a minimum instead, which is reported separately and excludes rather than penalises - and have a person make that exclusion, with a recorded reason.
Why is my leader only one point ahead?
Because the vendors are genuinely close. A gap of one or two points out of a hundred is smaller than the disagreement between two reasonable evaluators, so treat it as a tie and decide on something you can defend - references, capacity, risk - rather than on the decimal.
Are the scores stored or sent anywhere?
No. Supplier evaluations are commercially sensitive and sometimes legally disclosable, so everything runs in the browser with no upload and no storage. Download the CSV to attach to the decision note.
Last reviewed by the A2Z.Tools team against the sources listed above.