A claims handler settles forty claims this week. A solicitor makes twelve filings. A developer merges nine pull requests. Now rank them, and show the exchange rate.
There is no exchange rate, and the measurement literature has said so for decades. An evidence appraisal commissioned to find a universal measure of knowledge-worker productivity concluded that knowledge work is too varied and its outputs too intangible for a single universal measure to be possible.1 Not missing. Not premature. Not possible.
What fails has a name. Any sound productivity measure has to satisfy four conditions, and one of them is comparability: tasks and measurements should be the same across workers and over time.2 Cross-role output comparison fails that by construction. A settled claim and a merged branch are not the same task counted the same way.
So the question worth asking is not which output metric to adopt. It is which construct means the same thing in the claims team as it does in engineering.
A definitional problem, not a tooling gap
Economists got here decades before workforce analytics existed. Reviewing a multi-year project on service-sector output, Triplett and Bosworth reported no central theme to the problem of services measurement: every industry they examined contained unique problems, and measuring service output demanded a hedgerow-by-hedgerow assault rather than one method.3
One escape route is to narrow the scope, on the reasoning that a sector only needs a measure that works for itself. Comparability breaks there too. Financial-services output is captured indirectly, through Financial Intermediation Services Indirectly Measured and gross value added, and the fix proposed is different proxies by sub-sector — new loans per employee for banks, sales commissions per employee for insurers.4 One industry, two units.
If this were a tooling problem, an institution with every advantage would have solved it by now. The National Statistician's review of public services productivity records that as of June 2023 six of eleven service areas, defence and taxation administration among them, had no dependable inputs, no activity-based outputs and no quality adjustment. Where inputs are set equal to outputs, the review notes, the growth of productivity is by definition zero.5
Statutory data access, professional statisticians, a hundred-odd recommendations — and defence productivity is still a placeholder.
Nor does the survey fallback rescue it: subjective and objective productivity measures correlate poorly and cannot be used interchangeably.1 Activity and presence metrics, meanwhile, travel across roles for a reason that ought to disqualify them, which is that they measure nothing role-specific. Deloitte finds such metrics failing to account for invisible work, with 70% of surveyed workers saying their organisation already structures roles around problems to solve rather than repeatable tasks.6 None of that is laziness — it is structural. Measures of effectiveness are difficult to measure and sometimes impossible to define, so analysts settle for measures of performance that are easier to measure and easier to manipulate.7 Nobody chose keystroke counts out of malice. They chose them because the thing they actually wanted could not be defined.
The one construct that travels
Flow describes a state rather than a deliverable: a match between challenge and skill, sustained without interruption, with feedback tight enough to correct against.8 Nothing in that description changes when the output changes from a settled claim to a merged branch.
It holds inside a single job too. One developer's countable output swings from day to day for reasons that say little about how well they worked. Tuesday is nine merged pull requests. Thursday is a production incident that ends with nothing shipped and a real problem solved. Count deliverables and the week had one good day. Ask instead whether the work had that challenge-skill match and that uninterrupted stretch, and the question means the same thing on both days. Output is defined by the task; flow is indifferent to it.
In work settings it is also mostly a property of circumstance. An experience-sampling study found 74% of the variance in flow attributable to situational characteristics rather than dispositional ones.9 Were flow largely a trait, measuring it would be a personality test with extra steps. Because it is largely situational, it belongs to work design — and work design is something an operating model can actually move.
Then the number. The 500% productivity multiplier that circulates in flow writing, including WorkstyleIQ's own earlier research, should be retired. The best-evidenced estimate pools twenty articles containing twenty-two studies at r = 0.31, 95% CI [0.24, 0.38], and the same review states plainly that those studies are sport and computer-gaming tasks and that current evidence cannot determine the causal direction or the mediating mechanisms.10 A medium effect, out of domain, with the arrow pointing in an unknown direction.
That is a weaker headline and a better argument. A modest effect holding across two different task families is worth more to an operating model than a spectacular one that never left the anecdote that produced it. No output metric can be pooled at all.
Where it breaks, and what a board should ask
The most credible objection comes from inside the flow literature. The field is still arguing over a common conceptualisation, including whether to operationalise flow as continuous or discrete, and flow varies sharply between people and tasks.8 A construct whose own researchers cannot agree how to measure it is an odd thing to hand a board. Concede it: flow is contested at the level of operationalisation, output at the level of definition, which is the deeper failure. An organisation measuring flow imperfectly is at least measuring the same thing in every function.
The objection also shrinks with the unit of analysis. Whether flow can be measured consistently across an organisation is an open question. Whether one person's flow can be read against their own baseline is not: the individual is their own control, and a trend line needs no agreed unit to be legible. Someone watching their own flow across a quarter can see which days, which meeting loads and which rooms erode it, and change what is theirs to change. That signal is worth having before the organisational question is settled, and it costs nothing in comparability because it never leaves the person it describes.
A harder objection follows. The headline evidence is not about work at all; it comes from athletes and gamers, and causation is unresolved.10 An argument about transferability is resting its number on evidence that has not itself been shown to transfer. So claim less. The pooled figure is a floor on what a portable construct is worth, not a workplace uplift. A reader who finishes this believing flow causes a 31% productivity gain has been misled. The finding grounded in actual workplaces is the situational one, and it speaks to manageability rather than effect size.9
Whether any of this survives contact with a board pack turns on the last objection. The moment protected flow time becomes a target it turns into a measure of performance rather than effectiveness, easy to measure and easy to game,7 and relying on the wrong measure damages the productivity it was meant to raise.2 Which forces a narrow deployment: read flow in aggregate, use it to change how work is scheduled and interrupted, and never score it against a named individual or attach it to reward. The personal trend belongs to the person; the aggregate belongs to the operating model.
Tools built on row-level activity logs have no version of that restraint: the named individual is their unit of measurement by architecture. WorkstyleIQ reads workstyle patterns at the level of teams and operating rhythms, the only vantage point from which measuring the state rather than the output is an argument rather than a pitch.
None of this makes flow a solved measurement problem. It makes it the best available common denominator where the alternative is arithmetic on incommensurable units. Boards will go on asking what each function produced last quarter, and the answer will go on failing to compare with the one beside it. The better question is whether the conditions under which good work happens improved or degraded this quarter — and most boards, asked today, could not answer it.
Footnotes
-
Advanced Workplace Associates & Center for Evidence-Based Management. (2015). Measuring knowledge worker productivity (wpa INSIGHT). Workplace Performance Innovation Network. https://www.globalworkspace.org/wp-content/uploads/6FactorsKnowWorkerProductivity_wpaINSIGHT_Sept15.pdf ↩ ↩2
-
Sauermann, J. (2016). Performance measures and worker productivity. IZA World of Labor, 260. https://wol.iza.org/articles/performance-measures-and-worker-productivity/long ↩ ↩2
-
Triplett, J. E., & Bosworth, B. P. (2000). Productivity in the services sector. Brookings Institution. https://www.brookings.edu/articles/productivity-in-the-services-sector ↩
-
Economics Observatory. (2025). Why is it so difficult to measure productivity in the financial services sector? Economics Observatory. https://www.economicsobservatory.com/why-is-it-so-difficult-to-measure-productivity-in-the-financial-services-sector ↩
-
Office for National Statistics. (2025). National Statistician's independent review of the measurement of public services productivity. UK Statistics Authority. https://uksa.statisticsauthority.gov.uk/wp-content/uploads/2025/03/National-Statisticians-Independent-Review-of-the-Measurement-of-Public-Services-Productivity.pdf ↩
-
Cantrell, S., & Commisso, C. (2023). Why measuring productivity fails. Deloitte Insights. https://www.deloitte.com/us/en/insights/topics/talent/measuring-productivity.html ↩
-
Stumborg, M. F., Blasius, T. D., Full, S. J., & Hughes, C. A. (2022). Goodhart's law: Recognizing and mitigating the manipulation of measures in analysis. CNA. https://www.cna.org/analyses/2022/09/goodharts-law ↩ ↩2
-
Bartholomeyczik, K., Knierim, M. T., & Weinhardt, C. (2023). Fostering flow experiences at work: A framework and research agenda for developing flow interventions. Frontiers in Psychology, 14, 1143654. https://pmc.ncbi.nlm.nih.gov/articles/PMC10360049 ↩ ↩2
-
Fullagar, C. J., & Kelloway, E. K. (2009). Flow at work: An experience sampling approach. Journal of Occupational and Organizational Psychology, 82(3), 595–615. https://www.ovid.com/journals/joop/fulltext/10.1348/096317908x357903~flow-at-work-an-experience-sampling-approach ↩ ↩2
-
Harris, D. J., Allen, K. L., Vine, S. J., & Wilson, M. R. (2023). A systematic review and meta-analysis of the relationship between flow states and performance. International Review of Sport and Exercise Psychology, 16(1), 693–721. https://www.tandfonline.com/doi/full/10.1080/1750984X.2021.1929402 ↩ ↩2
