Task Exposure Index: Methodology
Version: v0.1, release v2026.Q3 Capability reference date: 15 September 2026 Built on: O*NET Database 31.0 Status: open for challenge. Every rating in this index is published with its reasoning and can be disputed at the level of a single work activity.
1. What this index measures
The Task Exposure Index measures, for each discrete task in the standard occupational taxonomy, how much of that task current AI systems can carry, and what stops them from carrying it, as of a stated calendar quarter.
It does not predict job losses. It does not forecast. It describes a capability frontier and the frictions standing in front of it, at a point in time, so that the same measurement can be repeated next quarter and the difference reported.
The index answers: "Of the work this job actually consists of, how much can a machine do right now, and where does that stop?"
2. Why the unit is the task, not the occupation
Existing AI-exposure measures score occupations. An occupation is not a thing a machine can or cannot do. It is a bundle of many things, some of which moved years ago and some of which will not move this decade. Scoring the bundle produces a number that is true of nobody's actual day.
O*NET, the US Department of Labor's public-domain occupational database, decomposes roughly 1,000 occupations into approximately 20,000 discrete task statements. Those tasks are linked to a controlled vocabulary of about 2,000 Detailed Work Activities (DWAs), which roll up through Intermediate Work Activities to 37 Generalized Work Activities.
We rate at the DWA level. Two reasons:
1. Reproducibility. Twenty thousand free-text task statements cannot be rated consistently by anyone. Two thousand standardised activity statements can, and the rating can be re-checked. 2. The mapping already exists. O*NET has already done the work of assigning tasks to activities. Every task inherits the ratings of the activities it maps to, by a published rule rather than a second act of judgment.
Occupation scores are then aggregated upward from tasks. No occupation is ever rated directly.
3. The two-factor structure
Most indices emit a single "exposure" number. That number silently conflates two different questions, which is why such indices are almost impossible to falsify. We separate them.
Factor C: Capability
Can a current, generally available AI system produce the work product of this activity?
Rated 0–4 against a capability reference fixed for the release quarter:
| Score | Meaning |
|---|---|
| 0 | Not at all. The system cannot attempt the work product. |
| 1 | Fragments only. A person still does the task; AI contributes pieces. |
| 2 | Draft quality. Output is a starting point requiring substantial rework. |
| 3 | Working quality. Output is usable after human review and sign-off. |
| 4 | End-to-end. Output is deliverable with spot checks only. |
Capability is rated against what is generally available in the quarter, not against research demonstrations or unreleased systems.
Factor F: Friction
What prevents the capable thing from actually being used?
Five sub-scores, each 0–3, summed to 0–15. These are the structural reasons capability does not become adoption. They are rated independently of C.
| Sub-score | Question |
|---|---|
| Embodiment | Does the activity require a body acting in physical space? |
| Presence | Must a human be the one present for the act to count, as in care, testimony, custody, negotiation or persuasion? |
| Accountability | Must a licensed or legally liable person sign the output? |
| Context | Does it require private, tacit, or real-time organisational knowledge a model cannot hold? |
| Verification cost | How expensive is checking the output, and how damaging is an error that goes undetected? |
Combining them
With c = C / 4 and f = F / 15, both on 0–1:
Exposure E = c × (1 − f) Assisted A = c × f Untouched U = 1 − c E + A + U = 1, by construction.
This says something deliberately strict: an activity is exposed only when AI can do it and nothing structural stands in the way. Where capability is high but friction is high too, the activity is assisted. A human stays in the loop, and the job changes shape without disappearing. Where capability is low, the work is untouched, whatever the friction.
The decomposition is the point. Any score can be opened up to show which of the six components produced it, which is what makes an individual rating arguable rather than oracular.
Banding
The three shares are the measurement. Bands are a reading aid on top of them, assigned by a rule stated here so anyone can reproduce it:
| Band | Rule | Meaning |
|---|---|---|
| Untouched | c < 0.5 | AI cannot produce the work product |
| Exposed | c >= 0.5 and f < 0.5 | It can, and little stands in the way |
| Assisted | c >= 0.5 and f >= 0.5 | It can, but a person stays in the loop |
4. How the ratings were produced
All 2,087 detailed work activities were rated against the rubric above. The first 150 were rated by the lead rater and became the calibration set. The remaining 1,937 were divided into seven contiguous slices and rated independently, each rater working from the same written brief, the same anchor examples, and no sight of the others' work.
Every rater additionally rated the same 40 activities drawn from the calibration set. Those 40 shared items are what the reliability figures below are computed from: eight raters, 40 activities, six dimensions, 1,920 paired judgments.
Inter-rater reliability
Agreement of each rater against the lead rater on the 40 shared activities:
| Dimension | Exact agreement | Within 1 point | Mean absolute deviation |
|---|---|---|---|
| Capability | 90.4% | 100% | 0.10 |
| Embodiment | 86.8% | 100% | 0.13 |
| Presence | 88.6% | 100% | 0.11 |
| Accountability | 81.4% | 100% | 0.19 |
| Context | 85.4% | 100% | 0.15 |
| Verification cost | 77.1% | 100% | 0.23 |
| All dimensions | 84.9% | 100% | 0.15 |
No rater showed meaningful systematic bias: the largest mean signed deviation from the lead rater on any dimension was 0.17 of a point.
Two things follow, and both are stated rather than buried. Verification cost is the least reliable dimension at 77.1% exact agreement, because it asks raters to judge how damaging an undetected error would be, which is the most contestable question in the rubric. No disagreement anywhere exceeded one point, which bounds how much rater choice can move a published score.
5. From activities to tasks to occupations
Task score. A task inherits the mean of its linked DWAs' component scores, then E, A and U are computed from those means. Tasks with no DWA linkage are excluded and reported as coverage loss, never silently imputed.
Occupation score. The importance-weighted mean of its task scores:
E_occ = Σ (E_task × w_task) / Σ w_task w_task = O*NET task importance × task relevance
Alongside the headline figure, every occupation page reports:
- the share of tasks in each band (exposed / assisted / untouched)
- the count of tasks scored and the count excluded for missing linkage
- the three tasks contributing most to the score, and the three contributing least
A single number that cannot be opened up is not evidence. Every number here can be.
6. Versioning and the time series
Every release is frozen and named: v2026.Q4, v2027.Q1, and so on. Each carries the capability reference date it was rated against.
Scores are never silently revised. A correction is issued as a numbered erratum against the release it affects. The current quarter's release is the live one; all previous releases stay published and addressable.
Each release publishes the delta against the previous quarter: which activities moved, by how much, and the stated reason for each movement.
This is the part of the index that cannot be copied after the fact. A competitor can replicate a snapshot in a week. Nobody can retroactively produce a quarterly series they did not run. The series accrues only to whoever starts it and keeps it.
7. What this index does not claim
Stated plainly, because an index that will not say what it cannot do should not be trusted about what it can:
- Exposure is not displacement. Whether exposed work becomes a lost job depends on
firm economics, labour law, union agreements, capital cycles and organisational inertia. We measure none of those.
- Friction ratings are judgments. They are argued, published per activity, and
open to dispute. They are not measurements and are never labelled as such.
- No employment or wage effects are modelled. Where wage and headcount figures
appear, they are sourced from public statistics and are context, not output.
- Capability ratings reflect general availability, not the frontier. A system that
exists but is not usable by an ordinary organisation does not move a score.
8. Provenance labels
Every figure published anywhere on the site carries one of three labels, and the label appears next to the number, not in a footnote:
| Label | Meaning |
|---|---|
sourced | Taken directly from a public dataset (O*NET, BLS), unmodified. |
computed | Derived from sourced inputs by a formula published here. |
judged | Our rating against the rubric above, with rater and date attached. |
No figure on this site is unlabelled. If one is, that is a bug worth reporting.
9. Falsifiability
Each quarterly release includes a small set of pre-registered predictions about what will move in the following two quarters, and scores the predictions made in the previous release, including the ones that were wrong.
An index that never says anything checkable is a branding exercise. This one keeps a public record of its own accuracy.
10. Sources
| Source | Used for | Licence |
|---|---|---|
| O*NET Database (US Dept of Labor / ETA) | Occupations, task statements, task ratings, work activities, work context, job zones | Public domain, attribution requested |
| BLS Occupational Employment and Wage Statistics | Wage and employment context | Public domain (US Government work) |
| BLS Employment Projections | Outlook context | Public domain (US Government work) |
O*NET® is a trademark of the US Department of Labor, Employment and Training Administration. This index is not endorsed by, affiliated with, or produced in cooperation with USDOL/ETA.
11. Changelog
| Version | Date | Change |
|---|---|---|
| v0.1 | 2026-09-15 | First release. 2,087 activities rated, 18,838 tasks and 923 occupations scored. Inter-rater reliability published. Two-factor structure (capability × friction), DWA-level rating, task-level publication, quarterly freeze with deltas, provenance labelling, pre-registered predictions. |