All articles
Future of Work9 September 2026·6 min read

Codified Knowledge Got Cheap. Tacit Knowledge Did Not. What That Changes About What You Assess

Stanford's payroll data shows a 19% employment gap for young workers in AI-exposed occupations, with declines concentrated in codified-knowledge roles while tacit-knowledge work grew. What the split does - and does not - mean for how you assess.

By AssessAll Editorial

Codified knowledge is what can be written down: procedures, formulas, documented facts, the contents of a manual. Tacit knowledge is what people acquire by doing and rarely write down — when to escalate, which exception is genuine, how to sequence a difficult conversation. AI has made the first kind cheap to retrieve. Most assessment budgets still measure it almost exclusively.

That mismatch is no longer a theoretical worry. It now shows up in payroll data.

What the 2026 payroll data actually shows

The Stanford Digital Economy Lab's Canaries in the Coal Mine? project tracks employment using ADP payroll records covering November 2022 through June 2026. Its August 2026 update reports that workers aged 22–25 in highly AI-exposed occupations have fallen roughly 19% behind their peers in less-exposed occupations — widening from about 15% a year earlier.

Three details matter more than the headline number:

  • There is no economy-wide displacement. The researchers state plainly that they do not observe widespread job losses attributable to AI. The effect is concentrated, not general.
  • The adjustment runs through hiring, not firing. Employment for young workers fell mainly because employers opened fewer entry roles, not because they separated more people.
  • The split is by knowledge type. Declines concentrated in roles built on codified knowledge. Employment among experienced workers grew in roles that lean on tacit knowledge built through experience and mentorship.

That third point is the one worth building policy around, and it is also the one to hold loosely. Occupational AI exposure is a proxy, not a measured causal channel, and the same lab has published a separate analysis of interest rates and timing as competing drivers of the same pattern. Treat it as a strong signal about what work is becoming scarce, not as proof of a mechanism. The live dashboard lets you check your own occupation groups rather than accepting the summary.

The distinction, stated precisely

The codified/tacit split is old — Michael Polanyi's "we can know more than we can tell" dates to 1966 — and in selection psychology it runs through Robert Sternberg and Richard Wagner's work on practical intelligence and tacit knowledge inventories in the 1980s and 1990s. Their claim was that a substantial portion of workplace competence is procedural, context-bound, and largely unspoken, and that it can be measured with scenarios rather than facts.

Worth knowing: the strong version of that claim has been seriously contested, notably by Linda Gottfredson, who argued the tacit-knowledge measures overlap heavily with general mental ability and that the incremental-validity claims were overstated. The honest position is narrower and still useful: scenario-based measurement of contextual judgment works empirically, whether or not "practical intelligence" is a distinct construct.

What the split looks like in a real role:

Codified, in a collections analyst role

  • The regulatory ageing buckets and provisioning rules
  • The scripted disclosure required at call opening
  • The system steps to raise a settlement request

Tacit, in the same role

  • Whether this particular customer's story warrants a restructure or a firm deadline
  • When a promise-to-pay is worth accepting and when it is a stall
  • How to escalate a wrong system recommendation without stalling the queue

The first list is a training manual. The second is what separates the analyst you keep from the one who churns at month nine.

This is not an argument that knowledge tests stopped working

It would be easy to over-read the data into "stop testing knowledge." The evidence says otherwise. In Sackett and colleagues' revised validity estimates, job knowledge tests carry an operational validity of about .40 — second only to structured interviews at .42, and above work sample tests at .33 and cognitive ability at .31. Knowledge testing remains one of the best-evidenced tools in selection.

Two different claims are in play, and conflating them produces bad decisions:

  1. A prediction claim: within a role, people who know more about the work perform better. This still holds, at the same strength as before.
  2. A demand claim: the number of jobs whose main output is retrieving and applying codified knowledge appears to be shrinking, and the value a new hire adds sits increasingly in judgment, exception handling and supervision of automated output.

Your assessment design has to answer the second claim without discarding the evidence behind the first. In practice that means keeping the knowledge component and adding a judgment component — not swapping one for the other.

Four ways to get judgment onto the score sheet

Scenario assessments with defensible scoring. Present a realistic dilemma with no clean answer and score the reasoning. The design trap is writing scenarios with an obvious "nice" option, which measures social desirability rather than judgment. AssessAll's AI-graded scenario items score open responses rather than multiple choice, which removes the option-elimination shortcut, and integrity bands from AI proctoring flag sessions where the response pattern does not match the candidate's behaviour.

Work samples built around incomplete information. A work sample that hands over every input tests execution. Deliberately omit one thing a real request would also omit and score whether the candidate notices, asks, or assumes. Assumption-flagging is a measurable behaviour.

Structured interviews that probe the decision, not the outcome. Behavioural questions reward polished stories. Follow-up probes — what did you consider and reject, what would have changed your mind, who did you consult — separate lived experience from a rehearsed narrative, and structure is what earns interviews their .42.

AI-supervision tasks. Give the candidate a plausible but subtly flawed AI output relevant to the role — a summary with an invented figure, code that passes tests but mishandles an edge case, a customer reply that is fluent and wrong — and score detection, correction and escalation. This is the closest thing to a direct measure of the skill the payroll data implies is now scarce, and it works precisely because AI assistance during the task is not something to defend against.

A seven-step rebalance

  1. Pick the three roles where your entry-level hiring has thinned the most in two years.
  2. For each, list the tasks a competent AI tool now does adequately. Anything on that list is a weak differentiator in screening.
  3. Write down what remains — decisions, exceptions, escalations, judgment calls under time pressure.
  4. Audit your current assessment: what percentage of scored points sits in the first list versus the second?
  5. Build two or three scenario or AI-supervision items from the second list, using incidents your own supervisors describe.
  6. Pilot them alongside your existing test on current employees whose performance you already know. Judgment items that do not separate strong from weak incumbents are not ready. Credit-based, pay-as-you-go assessment makes a pilot of a few dozen candidates cheap enough that this step actually happens.
  7. Re-check adverse impact after the rebalance, not before. New item types change the pass-rate pattern.

When codified knowledge should stay the main event

For licensed and regulated roles — clinical, aviation, legal, financial compliance, safety-critical operations — recall accuracy is the job requirement, and a rebalance toward judgment items would be both invalid and indefensible. Where a legal or regulatory standard defines minimum knowledge, that knowledge is the construct, and the EEOC's guidance on job-relatedness supports testing it directly. The same holds early in training pipelines, where a knowledge baseline is a prerequisite to acquiring judgment at all. The argument here is about the balance in general commercial roles, not about abandoning knowledge measurement.

The takeaway

The World Economic Forum's Future of Jobs Report 2025 projects that 39% of workers' core skills will change by 2030; the payroll data suggests the change is already sorting by knowledge type rather than by job title. If your assessment still allocates most of its scored points to what a manual contains, you are measuring the half of the job that got cheaper — and reporting a high score for it as if nothing had moved.

#tacit-knowledge#ai-and-work#future-of-work#work-samples#selection-science

Measure it, don't guess it.

Start free with 100 credits — or write to solutions@bodhih.com.

Start free