Calibration Is a Merge Sort
The room with the windows covered — how an organization turns many scores into one, and why it might be the high point of the whole job.
In review season, after assessment and after evaluation, comes a stage where the scores from many managers get reconciled and tuned against one another. It’s called _calibration._There’s no clean Korean translation, which tells you something; people reach for words like standardization or alignment.
Merge sort
At bottom, picture a higher-level manager doing a merge sort — a merge of sorted arrays — over the results their sub-managers handed up, against the higher manager’s own standard. That manager then acts as a single manager, and the whole thing recurses upward. Because you’re combining the evaluations of genuinely different teams, you spend the time persuading and defending: here’s my standard, here’s my evidence.
You try to align, but teams differ in their standards and their circumstances, so conflict is built in. Sometimes you have to decide which of two teams’ aces is the stronger — and so, now and then, an uber tech lead or some deliberately unrelated, objective party gets pulled in to offer a second read. You ask for extra feedback. You go find more objective evidence.
The conversation mostly centers on the outliers. One consequence is that it’s easy to coast quietly in the middle; another is that once you’re marked, for good or for ill, the mark travels far. You come away knowing the other teams’ aces and their people-to-watch, and when it’s your turn you’d better be ready to brag about your own. Goodwill banked with other teams tends to come back to you.
When calibration ends, the result starts carrying real weight. What began as your direct manager’s score becomes the organization’s score — “the calibration result,” “the org’s decision” — and the heavier label changes how it lands. At one company I mostly spent this time generating kind, constructive feedback to send back down. Though the direct manager, as ever, still owns the largest pile of homework.
The challenges
First, you must not see your own result in this process. By the same logic, a manager in the room shouldn’t see the results of peers at their own level. So you need conditional, gated access — and you get the small theater of managers stepping out, waiting in the hall, being called back in. The software built to support all this meets its hardest technical problem right here.
Get Chaesang Jung’s stories in your inbox
Then there’s the plain fact that everyone who needs to be in the room is busy, and getting them into one is brutal. Since you’re assembled anyway, you batch what you can, which means blocking at least half a day; with a deep org chart, you end up in several meetings. When the year gets planned, this is the first time you carve out, and people are gently asked not to take vacation across it.
Gathering physically is hard, and the security bar is high. It happens behind closed doors, sometimes with the windows covered. It’s one of the few defensible reasons for an international trip — and, using that as cover, teams sometimes decamp somewhere quiet off-site. At Google I was called in to these for a long stretch; the frequent reorgs meant the cast changed almost every time, and the tooling, as I understood it, was specialized precisely to make that handoff smooth.
Calibration in the AI era?
The process leans heavily on managers tuning one another’s judgment, which makes it most useful at large scale. Whether it’s needed in the AI era, or in a small org, is fair to debate — and the argument that the whole apparatus is unnecessary is one I understand. The catch is the side effect: without it, you get overridden by the instinct or the taste of one or two executives.
Handing the feedback and the results to AI to write is understandable, as conveniences go — but this is a place where people live, and it shows. In Korea especially, where engineering managers are much harder to come by than in the US, building a system that actually fits the org is far harder. AI trained on American norms, or the communities and YouTube videos that teach you from some other company’s story, sit a long way from the reality in front of you. So checking who’s doing well matters more, not less — it’s how you decide who to hand the bigger work to — and where people would rather stay invisible and middling, that’s exactly where the friction shows up.
Looking back over a long run of these, there were good scores and bad ones. From the receiving side, the cycles that followed a detailed, positive, constructive review were the good ones. The best message I can remember ran roughly like this: you solved the hard problem, your leadership is strong too — what we need now is bigger impact, so next cycle, take this project and drive it all the way through. Back in my day. I get a little sentimental.
ps. Before I ever joined Google, “calibration” meant something else entirely to me — _tslib_, _ts_calibrate_, the little utility for a touchscreen. How long ago was that?
Part of Lessons from the Company — first-person notes from inside engineering organizations, from my years at Google and after. See the full series →
Adapted from my Korean essay on Brunch: brunch.co.kr/@chaesang/169
Comments