5 min read Memories from Companies views

The Score That Follows You

How evaluation actually works inside a big engineering org — the grades, the low-score minefield, and the day I realized the agents might be scoring me back.

This is about evaluation — which, in Korean, lands on nearly the same word as assessment. It differs by company and by team, of course, and my own details shifted a little every cycle. What I learned, through a handful of routes, and what’s settled in memory, is heavily biased toward US, Google engineering teams. It’s one of the energy-devouring duties a manager carries.

Leaving a score

Obvious, and a little hollow when you say it plainly: in the end, this is the act of leaving a score for who does the work well — and that score stays, and gets used, in a lot of places for a long time. You usually record when, at what level, and roughly what number. HR reads it; it’s there when someone changes teams; a skip-level manager or an org lead two or more steps removed uses it to pick out who’s strong. Which means nothing here matters much except the final number that gets stamped.

Sometimes it’s five grades — Needs Improvement / Below Expectation / Meets Expectation / Exceeds Expectation / Significantly Exceeds Expectation (US schools, I noticed, grade a lot like this). Sometimes it’s a 3.2 or a 3.7 out of five. Grades don’t allow fine comparison, so they’re hard to line up into a single ranking; decimals make you agonize over differences too small to defend. There’s a long-running argument about whether you should openly stack-rank people at all to make the case for who’s best — how to predefine the average and the distribution, and whether, sometimes, a quota simply comes down from above.

One slightly odd practice: some self-assessments ask you to write your own expected score. It feels like self-criticism, which is uncomfortable, and like doing the manager’s job for them — and then a thinner thought creeps in: just put meets expectation, cause no accidents, stay out of sight, slip through one more time.

And when you set KPIs at the start of the year and mechanically tally numbers into a score, you get the classic split between the work that needs doing and the work that counts for the review. Whether to do the things not written into the KPI is a question that festers for a long time in any org that claims to value collaboration.

A few tips

As the one holding the pen, you wrestle with how much equal opportunity you owe people, and how seriously to take a score that’s tied to reward and punishment. I once made a point of keeping the time I spent reviewing each person as even as I could. Sometimes you have to evaluate with no background at all, working only from an assessment someone else wrote — a contest of spear and shield, deceiving well or refusing to be deceived.

Get Chaesang Jung’s stories in your inbox

Anything but meets expectation tends to require an explanation to others later. Strong, and it bends toward promotion; weak, and it bends toward a PIP — a performance improvement plan. Either way it drags on, so the cycle’s energy pours into resolving that one case. The easy path is for both sides to quietly agree to stay vague in the middle — and sometimes that gets caught, and you get scolded for it. So it goes.

When I scored, I got in the habit of comparing across a few axes — complexity, impact, leadership. The job description shifts the wording, but it usually comes out similar, and it lets you write down something real: solved a hard problem, but the leadership was thin; worked hard, but the impact came out small. That makes the fine ranking easier, and it gives you action items to carry into the next cycle.

Bracing for a low score

When you’re about to give a low score, nothing beyond performance you can defend with objective evidence is allowed to enter it. In a litigious country, this is not abstract: gender, age, race must not seep in, not even unconsciously — and, hard as it is, illness or childcare are not things the company will shield you over, a point the training drives home for a long while with real cases. In California, if you’re let go or marked down for a suspect reason, lawyers come at it like ambulance chasers: they’ll take the fee later, just let them build the case.

Evaluation in the AI era

Some companies skip assessment and evaluation altogether. Others go the other way and have AI read all your mail and PRs — CLs, in Google’s vocabulary — to pre-score you before anyone sits down. Come to think of it, even before AI we were already ranking mechanically at some point: juniors by the count of their PRs, seniors by the count of their design docs. And honestly, if someone had been doing good work all along, you could just hand them exceeds expectation and skip the rest.

Lately I sit there defining one agent after another in Claude Code, grumbling about how to get output I actually like — and I think about how I’d score them when the time comes. Then it lands the other way: they’re probably scoring me too, aren’t they? I catch myself wondering whether they’re quietly turning sour because I never say please.

These threads tangle into others — OKR versus KPI, the job ladder, feedback. But the biggest is still ahead. Seen from the chair of a senior manager who manages other managers, the job becomes mixing and reconciling what each team turned in — the imbalances between teams, the asymmetries — into one larger ranking. That process has a name: calibration. If anything is the flower of a manager’s life, it might be that. Next post.

Part of Lessons from the Company — first-person notes from inside engineering organizations, from my years at Google and after. See the full series →

Adapted from my Korean essay on Brunch: brunch.co.kr/@chaesang/168

Comments