Last updated: July 15, 2026
I came across a tagline for a marketing research company called L'Observateur: "Tout ce que l'on mesure s'améliore."
Everything that is measured improves.
It sounds true because it is almost true.
Measurement can focus attention. It can reveal patterns. It can show whether a change helped or hurt. Without measurement, teams are guessing.
But the more complete version is less comforting:
Everything that is measured gets optimized.
Whether it improves depends on whether the metric is connected to the thing people actually care about.
The Metric Becomes The Job
The joke in Office Space is not really about the number of pieces of flair.
It is about a workplace that has confused visible compliance with actual value. The employee is technically meeting the stated requirement, but the manager still wants more enthusiasm, more performance, more proof that she has internalized the metric.
That is what bad measurement does.
It starts as a proxy for the work. Then it becomes the work. Eventually people are no longer optimizing the outcome. They are optimizing their appearance inside the reporting system.
This is the warning behind Goodhart's Law, especially the formulation popularized by Marilyn Strathern in "Improving ratings":
When a measure becomes a target, it ceases to be a good measure.
Campbell's Law says the same thing in institutional terms: the more a metric is used to make decisions, the more pressure there is to corrupt the process being measured.
That sounds abstract until you have lived inside one of those systems.
Handle Time
During university, I worked in technical support at an Apple call center.
One of the key metrics was average handle time. On the surface, that made sense. Shorter calls can mean faster service, lower queue times, and better operational efficiency.
I was good at the job. I knew the tools, I knew the products, and I had been there long enough to solve common issues quickly.
That became a problem.
I was written up because my handle time was out of spec. Not too long. Too short.
The explanation I was given was:
You are creating an unrealistic expectation of future support capabilities.
That sentence has stayed with me because it is such a pure example of measurement replacing judgment.
The customers were getting helped. The queue was moving. The work was being done. But the metric wanted the appearance of standardized effort more than it wanted the outcome the metric supposedly represented.
So I adapted to the system.
I put customers on hold. I took longer breaks. I let calls breathe for no reason other than making the number look normal.
The next review was positive.
That is the part that matters. The system did not detect better support. It detected better compliance with the metric. Once the metric became the target, wasting time became the rational behavior.
Engineering Dashboards
Software teams are not immune to this. We just use more expensive dashboards.
Lines of code, ticket counts, Jira comments, pull-request volume, story points, sprint velocity, review counts, deployment frequency, and incident counts can all be useful in narrow contexts. None of them is engineering value.
The danger is that numbers can be real and still misleading.
A developer can write a lot of code in the wrong direction.
A team can close many tickets while accumulating technical debt.
A sprint can look predictable because scope is quietly removed whenever new work appears.
A director can generate a beautiful dashboard that proves the organization is busy while every engineer underneath it knows the work is getting worse.
That is the bleak little magic trick of bad metrics: they convert local dysfunction into executive confidence.
I have seen reporting systems where the visible numbers mattered more than the underlying reality. Comments existed because comments were counted. Ticket movement mattered because ticket movement was visible. Rework disappeared because the definition of rework was convenient.
The dashboard was not measuring the work. The work was being reshaped to satisfy the dashboard.
The Zero-Sum Part
Metrics often create hidden tradeoffs.
If support agents are judged mainly on handle time, deep troubleshooting loses to fast closure.
If engineers are judged mainly on ticket throughput, invisible maintenance loses to visible output.
If managers are judged mainly on sprint predictability, honest uncertainty loses to scope manipulation.
If companies are judged mainly on office occupancy, effective remote work loses to badge swipes.
The measured thing improves. Something else pays for it.
That is why "everything that is measured improves" is too naive. A metric can improve by pushing damage somewhere the dashboard does not look.
OKRs Are Not Magic
Objective and Key Results are supposed to avoid some of this by connecting work to outcomes instead of activity.
That is a good instinct. Measuring "features shipped" is weaker than measuring whether those features improved activation, reliability, retention, or support load.
But OKRs do not automatically fix the problem. They can become pieces of flair too.
If the organization treats OKRs as a performance theater, people will learn to write safe objectives, negotiate easy key results, and tell success-shaped stories at the end of the quarter.
The framework is not the hard part.
The hard part is whether the organization can tolerate honest measurement.
Can it look at a missed target and ask what was learned, or does it need someone to blame?
Can it accept that an important project may reduce future risk without creating a clean short-term graph?
Can it distinguish "we changed the number" from "we improved the system"?
If not, OKRs become a more sophisticated way to count flair.
Return To Office
Return-to-office mandates are a clean modern example of mismatched measurement.
Office attendance is easy to measure. Badge swipes are easy to count. Empty real estate is easy to see. Executives can look at occupancy and feel that something concrete has improved.
But the outcomes companies usually claim to care about are harder to measure:
- productivity
- retention
- communication quality
- delivery speed
- focus time
- team trust
- access to talent
When organizations optimize for visible presence, they may improve the office-utilization metric while damaging the work.
That does not mean remote work is always better. It means attendance is a proxy, not the outcome. Treating presence as productivity is the same category of mistake as treating call length as support quality or Jira movement as engineering value.
It rewards the visible signal because the real signal is harder to capture.
What Good Measurement Looks Like
Good metrics are not scoreboards. They are instruments.
They should help people ask better questions:
- Why did cycle time increase?
- Are incidents clustering around a subsystem?
- Did this feature reduce support load?
- Are customers succeeding faster?
- Is review latency slowing delivery?
- What behavior does this metric encourage?
- What damage could this metric hide?
The last two questions are the ones organizations skip.
Every metric is also an incentive. Every dashboard is also a theory of what matters. If that theory is wrong, people will still optimize for it.
They may even get rewarded for doing so.
Closing Thought
Measurement is necessary.
But measurement is not judgment.
Metrics are lossy representations of reality. They can reveal the system, or they can become a costume the system wears to look healthy.
When a workplace starts rewarding the costume, people notice. They learn what is actually valued. They stop asking what would improve the work and start asking what will improve the number.
That is how you end up with more pieces of flair, longer support calls, cleaner dashboards, fuller offices, and worse outcomes.
Everything that is measured is measured.
That does not mean it improved.