The Problem With the Numbers Everyone Is Already Citing
Organizations are making increasingly consequential decisions about artificial intelligence (AI). They are investing in systems intended to improve productivity, accelerate decision making, automate tasks, reshape workflows, and augment human performance. The capital allocation decisions are being largely based on artificial intelligence return on investment (AI ROI) figures that would not survive methodological scrutiny in any other domain of applied psychology. The pattern is almost always the same: A metric moves after a deployment, the movement is attributed entirely to the deployment, and the resulting ratio is presented to a board as fact. No comparison condition. No accounting for concurrent initiatives, seasonality, or regression to the mean. No estimate of what would have happened anyway. The difficulty is that movement is not necessarily impact.
Industrial-organizational (I-O) psychology solved this exact problem for a different class of intervention decades ago and built a well-established tool to solve it: utility analysis. Most practitioners know the name. Fewer use it regularly, and even fewer have seen it applied outside its original context of employee selection. Its premise is based on a question raised in I-O psychology that “When an organization invests in an intervention intended to improve performance, how much economic value does that intervention actually create?” The objective was to answer whether a selection procedure, training program, or other human capital intervention was actually worth its cost. This article proposes that the logic of utility analysis may provide a useful starting point for thinking about a new organizational phenomenon: artificial capital (AC)—the productive capacity embodied in an AI system and deployed into organizational workflows. The purpose here is not to present a validated artificial capital utility (ACU) model. It is to propose a conceptual extension and identify the methodological questions that I-O psychologists could help answer.
What Is Utility Analysis? A Practitioner’s Primer
Utility analysis can appear mathematically intimidating, but its underlying question is remarkably practical. Was this intervention actually worth what it cost? It was originally developed to evaluate whether a new employee selection test, or a training program, produced enough improvement in job performance to justify its price tag (Brogden, 1949; Cronbach & Gleser, 1965). Before utility analysis, organizations often adopted new selection tools because they seemed reasonable or because a vendor recommended them—with no dollar-value estimate of whether the tool actually paid for itself.
Imagine an organization is considering a selection procedure. The organization wants to know: If we use this procedure to make better hiring decisions, what is the expected economic value of doing so?
Traditional utility analysis attempts to answer that question by considering five key factors, each answering an intuitive question:
- N—How many people are affected?
A useful intervention may have modest value if it affects only a handful of positions and much greater aggregate value if it affects hundreds or thousands.
- T—How long does the benefit last?
An intervention that produces benefits over several years has a different economic value from one whose effects disappear quickly. - SDy—How economically valuable is performance variability?
Not all performance differences have the same economic significance. The value associated with performance variability depends on the role and the economic consequences of performance in that role.
- dt—How much does the intervention actually affect performance?
This is the critical effectiveness question. Does the intervention meaningfully change performance?
- C—What does the intervention cost?
The economic value of an intervention must ultimately be considered net of its implementation costs.
Put together, the classical formula is: ΔU = N × T × SDy × dt − C
Schmidt and Hunter (1983) and Schmidt et al. (1979) extended and defended this model through decades of application to selection systems specifically, establishing both its practical utility and the boundaries of its assumptions. The model’s structure does not assume anything about the nature of the intervention—only that it is something applied to a defined population, over a defined period, with a measurable effect on performance that varies in economically meaningful ways. Nothing in that structure requires the intervention to be human delivered.
The Proposed Extension: From Human Capital Intervention to Artificial Capital
Artificial capital, as used here, means the productive capacity an organization gains by deploying an AI system into a workflow—analogous to how “human capital” refers to the productive capacity a person brings, not the person themselves.
The classical utility model was built around interventions applied to individuals—a test taken once, a training program completed once. Nothing in the model’s underlying logic actually requires a human intervention, only that something measurable is applied to a defined group of people, over a defined time, with a measurable effect. That means the same structure can, in principle, be applied to an AI deployment: ΔU = N × T × SDy × dt – C
Here, N becomes the number of roles the AI deployment touches, T becomes how long the deployment is expected to remain useful, SD_y stays exactly the same (it’s a property of the job, not of what’s improving performance in it), and d_t (AC) — the new term becomes the measured effect of the AI deployment on actual performance in those roles.
To be clear about what’s established and what’s new: the formula’s structure, and every term except d_t(AC), comes directly from decades of existing utility analysis research. The proposed extension is narrower than it might first appear – it is the claim that this structure can be validly applied to a non-human intervention, and the practical work of figuring out how to estimate d_t (AC) in that new context.
A Worked Example
Consider a midsized professional services firm deploying an AI research assistant across its analyst team of 40 people, and expecting it to remain useful for roughly 3 years before requiring significant retraining or replacement.
The firm estimates SD_y conservatively at $40,000, a standard approach is to use roughly 40% of average salary as a proxy for the dollar value of a one standard deviation difference in performance, and this firm’s analysts earn approximately $100,000 on average.
To estimate d_t (AC), the firm compares output quality and speed between a group of analysts using the tool and a matched group not yet using it, over a 3-month pilot. The comparison shows a standardized effect of 0.5, a moderate, believable improvement, not an inflated marketing claim.
Implementation cost, including licensing, integration, and training, comes to $150,000.
ACU = 40 x 3 x $40,000 x 0.5- $150,000 = $2,250,000
Compare this to what the firm’s initial vendor pitch claimed: a flat “35% productivity increase,” with no comparison group, no cost accounting, and no way to verify the number. The utility analysis estimate is more conservative, but it is also defensible in a way the vendor’s figure never was—every input can be examined, challenged, and re-estimated with better data as they becomes available.
What Actually Changes: The Nonhuman Intervention Problem
The substitution above is not merely notational. Three structural differences deserve direct treatment, not silent assumption.
First, classical utility analysis assumes a discrete, bounded intervention applied to individuals—a test administered once, a training program completed once. Artificial capital is continuously present, frequently updated, and its “dose” is not uniform across the population it touches; a given system may be more heavily used, or used differently, by different role holders. The T term therefore needs to account not just for duration but for a capability that may itself change materially within the measurement window—something classical utility analysis, applied to a static selection instrument, never had to model.
Second, classical utility analysis assumes the intervention affects individual performance, aggregated to the population level. Artificial capital frequently operates at the team or workflow level, restructuring how work is distributed among people rather than simply improving each person’s independent output. This argues for a more careful specification of the unit of analysis than the classical model requires, and possibly for a team-level extension of SDy that the existing literature has not needed to develop.
Third, and most consequentially: classical utility analysis has never had to account for an intervention’s own depreciation during the measurement period. A selection test’s validity is treated as stable across the analysis window. Artificial capital’s effective capability can degrade—through model drift, changing task demands, or simple obsolescence relative to newer alternatives—on a timescale far shorter than the deployment’s nominal useful life. Any credible application of this model to artificial capital needs an explicit obsolescence adjustment that the classical formula does not carry.
The Central Estimation Challenge: Populating dt (AC)
This is where honesty matters most, and where I would ask reviewers and practitioners alike to hold the field to real standards rather than accept convenient shortcuts. Selection utility analysis has the benefit of nearly a century of accumulated meta-analytic validity coefficients (Schmidt & Hunter, 1998) to draw on when estimating dt. Artificial capital deployments have no comparable accumulated base yet. Three estimation routes are available now, in descending order of rigor and ascending order of practical feasibility:
- Controlled pre/post rollout, comparing matched groups with and without the deployment, ideally rolled out in stages—the approach used in the worked example above, and the strongest available option today.
- Matched-team comparison, where similar teams that have and have not yet received the deployment are compared at the same point in time.
- Structured supervisor and output-based rating differentials, the least rigorous of the three but the most immediately deployable, and meaningfully better than the undifferentiated before/after correlation currently dominating industry AI-ROI claims.
None of these fully replicates the accumulated evidentiary base selection utility analysis enjoys. All three are a genuine improvement over what organizations are currently doing, which in most cases is nothing resembling controlled comparison at all.
Old Critiques, Applied to New Terrain
Utility analysis has never been without its critics, and those critiques do not disappear when the intervention changes. Cascio’s (1980) long-standing concerns about the volatility and defensibility of SDy estimates apply with undiminished force here—arguably more force, because organizations attempting to estimate performance variability in roles newly reshaped by AI deployment have even less stable a baseline to work from than the roles selection researchers originally studied. Boudreau’s (1983) refinements to the classical model, accounting for variable costs and taxation effects on utility estimates, are directly relevant to any serious organizational application of ACU and should be incorporated rather than treated as optional.
The honest position is not that these critiques are resolved by moving the model to a new domain. It is that the model, imperfect as it has always been, remains more defensible than the alternative currently in use across most organizations. The arrival of artificial capital gives I-O psychology a new context in which its existing measurement discipline may be useful—and new methodological problems that the field will need to solve.
What This Looks Like From Inside an Organization
Practitioners evaluating an AI deployment are rarely asked “is this statistically defensible.” They are asked “was this worth it,” under time pressure, with incomplete data, by leaders who have already decided the answer is yes. The value of a framework like this is not that it produces a single, unimpeachable number. It is that it forces the right questions into the room before the number gets presented to a board: What population, specifically, did this touch? Over what window are we actually measuring? What would SDy conservatively be for these roles if we’re honest about it? And critically—what’s our actual comparison condition, or do we not have one? Organizations that cannot answer that last question are not measuring utility. They are measuring a coincidence in time.
Why This Matters to I-O Psychology?
This development also raises a broader question about the role of I-O psychology.
Organizations are increasingly asking:
- Which AI systems should we deploy?
- What work should be automated?
- What should remain human?
- How should jobs be redesigned?
- How does AI affect performance?
- What happens to productivity?
- How should organizations evaluate the return on AI-enabled work?
These questions sit across multiple disciplines.
Computer science can help determine what systems can do.
Finance can help evaluate investment.
Economics can help examine productivity and market effects.
But I-O psychology brings a particular body of expertise to the problem: measurement, performance, individual differences, work design, organizational interventions and the relationship between human capability and organizational outcomes.
Artificial capital therefore presents an opportunity for the field to contribute to a conversation that is already underway.
The opportunity is not to claim ownership of AI economics. It is to bring I-O’s measurement discipline into the emerging economics of AI-enabled work.
A Call for a Research Agenda, Not a Finished Instrument
The proposed extension should be treated as a research agenda rather than a validated instrument. What the field needs now, and what I would welcome, is collaboration on rather than claim to have solved alone: accumulated case data on dt (AC) across deployment types, to begin building the kind of meta-analytic base selection research has enjoyed for decades; a formal treatment of the team-level unit-of-analysis question; and an honest empirical test of whether obsolescence rates for AI capability follow any regularity stable enough to model, or whether they remain genuinely deployment specific.
I-O psychology has spent a century building the most rigorous available methodology for pricing the value of interventions applied to people. Artificial capital is not people, but it is squarely the kind of organizational intervention this field’s oldest tools were built to evaluate.
The research agenda can be expressed through several questions that call for empirical attention.
- Estimating (d_t(AC)): Can sufficiently comparable deployment data eventually produce reliable estimates of AI intervention effects across different work contexts?
- The unit of analysis: When AI changes workflows rather than merely individual performance, how should utility analysis model individual, team and workflow effects?
- Artificial capital depreciation: Do AI capabilities depreciate according to sufficiently regular patterns to be incorporated into a utility model, or will depreciation remain highly deployment specific?
- Measuring the dose: How should we measure actual exposure to artificial capital when employees use the same system differently?
- Human–artificial interaction: What portion of observed performance improvement is attributable to the AI capability itself, and what portion results from changes in human behavior, workflow design, learning, or adaptation?
- Longitudinal utility: How should organizations evaluate an AI intervention whose capability, users and surrounding workflow are all changing simultaneously?
Conclusion: A New Intervention Enters the I-O Conversation
Utility analysis emerged from a practical organizational question: Does an intervention create enough value to justify its cost? That question still requires an answer even though the intervention has changed from human capital to artificial capital.
Artificial capital introduces productive technological capability directly into organizational workflows. Unlike many traditional interventions, it may be continuously available, unevenly used, embedded in team-level work, and subject to rapid capability change. Those characteristics create methodological challenges that cannot simply be assumed away.
But they also create an opportunity. I-O psychology has a long intellectual tradition of asking organizations to move beyond intuition and anecdote toward defensible measurement of human and organizational performance.
Artificial capital invites that tradition into new territory. The proposition advanced here is deliberately modest: Artificial capital may be evaluated as an organizational intervention using an extension of the utility-analysis logic—but that extension requires empirical development before it can be treated as an established model.
The immediate task is therefore not to produce another attractive AI ROI calculator. It is to build the evidence base, to identify the appropriate units of analysis, to establish credible comparison conditions, to understand capability depreciation, to develop defensible estimates of intervention effects, and ultimately to determine whether artificial capital utility can become a meaningful addition to the I-O toolkit.
The question for the field is not whether AI will enter organizational work. It already has. The question is whether I-O psychology will help organizations measure what happens next.
References
Boudreau, J. W. (1983). Effects of employee flows on utility analysis of human resource productivity improvement programs. Journal of Applied Psychology, 68(3), 396–406.
Brogden, H. E. (1949). When testing pays off. Personnel Psychology, 2(2), 171–183.
Cascio, W. F. (1980). Costing human resources: The financial impact of behavior in organizations. Reston Publishing.
Cronbach, L. J., & Gleser, G. C. (1965). Psychological tests and personnel decisions (2nd ed.). University of Illinois Press.
Schmidt, F. L., & Hunter, J. E. (1983). Individual differences in productivity: An empirical test of estimates derived from studies of selection procedure utility. Journal of Applied Psychology, 68(3), 407–414.
Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274.
Schmidt, F. L., Hunter, J. E., McKenzie, R. C., & Muldrow, T. W. (1979). Impact of valid selection procedures on work-force productivity. Journal of Applied Psychology, 64(6), 609–626.
Wonder Jonamu is a registered industrial and organizational psychologist and founder of People Capabilities Advisory Group (PCAG). Correspondence regarding this proposed extension is welcome- wonder@peoplecapabilities.com, wonderjonamu@gmail.com
Volume
64
Number
2
Author
Wonder Jonamu, People Capabilities Advisory Group (PCAG)
Topic
Business, Industrial & Organizational Psychology