Rating Distribution Analyzer
Performance rating spread by manager, with small teams held to the standard they deserve instead of topping the list by accident.
Reads .xlsx and .csv inside your browser. Nothing is uploaded.
Nothing here is evidence of bias
A manager whose team genuinely does better should rate them better, and a team that had a hard year should show it. What this tool produces is a shortlist of distributions worth asking about. The asking is the point. A number on this page is the beginning of a conversation with a manager, never the end of one.
The review export
Drop a spreadsheet here
One row per person, with their rating and who rated them.
Sorting managers by percentage mostly sorts them by team size
A manager of four who gave three top ratings is at 75%. In an organization running at 20% that looks damning. It is also roughly what chance does with four people often enough that saying it out loud would be unfair to them.
Meanwhile a manager of forty at 60% is a far stronger signal, and on a list sorted by percentage they sit below the manager of four. Both ends of that list fill up with small teams, because small teams vary more. The conversation that follows lands on the wrong people, and the manager of four has to defend an artifact.
Each group here is tested against the organization’s own rate with an exact binomial test, so a small team has to be far more extreme before anything is said about it. The list is ordered by average position on the scale rather than by share at the top, because ordering by share would undo the whole point in the interface.
You did not test one manager, you tested all of them
Look at fifty managers at a one in twenty threshold and two or three will look unusual, even if every manager in the building rates identically. Almost nothing corrects for this, which is how a statistical artifact becomes a meeting.
Findings come in two tiers. A group that stands out survived a correction for having looked at everybody at once. A group worth a look did not, and some of those will be nothing. The number expected by chance is printed beside the number found, so the two can be compared by eye.
Everyone getting the same rating is the finding no test catches
A statistical test compares how often something happens. It cannot see a manager who gave every single person the same rating, because there is nothing unusual about the rate. That is called out separately, and it is usually worth more attention than anything a test does find. A scale where almost nobody is distinguished from anybody is not doing the job it was built for, whatever the ratings are being used to decide.
The order of the scale is the first thing to check
A spreadsheet does not say which end is good. Sorted alphabetically, “Exceeds” comes before “Meets”, which would invert every finding on the page while looking perfectly reasonable. Numbers order themselves and common wording is recognized, including the difference between “meets” and “does not meet”. Anything else is handed back as a guess for you to fix before reading further.
Position on a scale is not a mean rating
The distance from the bottom of a rating scale to the middle is not known to equal the distance from the middle to the top. Averaging the numbers invents an arithmetic the scale does not have. Groups are ordered by average position, which is used for sorting and never reported as a score.