The Task Exposure Indexv2026.Q3
Occupation · SOC 15-2051.00 · Job Zone 4

AI exposure: Data Scientists

Develop and implement a set of techniques or analytics applications to transform raw data into meaningful information using data-oriented programming languages and visualization software. Apply data mining, data modeling, natural language processing, and machine learning to extract and analyze information from large structured and unstructured datasets. Visualize, interpret, and report data findings. May create dynamic data reports.

Reading this score

computed

Data Scientists sits in the top 1% of every occupation measured. 69.1% of what this job consists of, weighted by how important each task is to the role, is work current AI systems can produce with little standing in the way. Very few occupations score this high. The ones that do tend to share a trait: the output is a document, a calculation or a message, and nobody has to be in a particular room for it to count.

What holds the line here is context. Across this occupation's 16 tasks it averages 1.62 out of 3, the highest of the five friction dimensions. In plain terms, the work depends on knowledge the model cannot hold. Much of this job runs on things that were never written down: what this particular organisation does, what happened last week, what the person across the table actually meant. That context is the barrier, and it erodes as systems are given more access.

The most exposed thing this job does is Clean and manipulate raw data using statistical software, at 86.7%. The least is Deliver oral or written presentations of the results of mathematical modeling and data analysis..., at 26.7%. A gap of 60.0% between two parts of the same job is the reason this index publishes at task level. An occupation-wide number would have hidden both.

Within computer and mathematical occupations, this one is more exposed than most. The median across the 36 roles in the group is 56.7%, and only 2 of them score higher than this. Occupational families are not uniform, and the spread inside them is often wider than the gap between them.

What would move this score. Of 16 tasks, 16 are currently banded exposed, 0 assisted and 0 untouched. For that distribution to shift materially would take systems being given deeper access to the organisation's own records and history, which is already happening. The score is re-computed every quarter against a fresh capability reference, and the change is published rather than quietly applied.

Task by task

16 tasks, O*NET 31.0
TaskExposedAssistedUntouchedImportanceBand
Clean and manipulate raw data using statistical software.86.7%13.3%0.0%4.04exposed
Identify relationships and trends or any factors that could affect the results of research.86.7%13.3%0.0%3.67exposed
Recommend data-driven solutions to key stakeholders.80.0%20.0%0.0%4.17exposed
Compare models using statistical performance metrics, such as loss functions or proportion of explained variance.80.0%20.0%0.0%4.04exposed
Identify solutions to business problems, such as budgeting, staffing, and marketing decisions, using the results of data analysis.80.0%20.0%0.0%3.91exposed
Propose solutions in engineering, the sciences, and other fields using mathematical theories and techniques.80.0%20.0%0.0%3.05exposed
Create graphs, charts, or other visualizations to convey the results of data analysis using specialized software.76.7%23.3%0.0%4.33exposed
Analyze, manipulate, or process large sets of data using statistical software.73.3%26.7%0.0%4.33exposed
Apply feature selection algorithms to models predicting outcomes of interest, such as sales, attrition, and healthcare use.73.3%26.7%0.0%3.75exposed
Write new functions or applications in programming languages to conduct analyses.73.3%26.7%0.0%3.61exposed
Apply sampling techniques to determine groups to be surveyed or use complete enumeration methods.66.7%33.3%0.0%3.00exposed
Read scientific articles, conference papers, or other sources of research to identify emerging analytic trends and technologies.65.0%10.0%25.0%3.38exposed
Identify business problems or management objectives that can be addressed through data analysis.55.0%20.0%25.0%4.13exposed
Design surveys, opinion polls, or other instruments to collect data.55.0%20.0%25.0%2.95exposed
Test, validate, and reformulate models to ensure accurate prediction of outcomes of interest.50.0%25.0%25.0%4.25exposed
Deliver oral or written presentations of the results of mathematical modeling and data analysis to management or other end users.26.7%23.3%50.0%4.21exposed

Task text and importance ratings sourced from O*NET 31.0. Shares computed. The occupation score is the importance-weighted mean.

Where the score comes from

judged

Every task is scored through the standardised work activities it maps to. These are this occupation’s averages on the six rubric dimensions. Capability is what AI can do; the other five are what stands in the way.

DimensionMeanScale
Capability3.620-4
Embodiment0.060-3
Presence0.120-3
Accountability0.530-3
Context1.620-3
Verification cost1.310-3

What this means in practice

Where most of a role's weighted task load is exposed, the work that survives is usually the part of the job nobody wrote into the job description: deciding what should be produced rather than producing it, and being answerable for the result. The tasks lowest on this page are a better guide to where to spend your time than any general advice about the future of work.

Occupations either side of this one

The four closest scores in the same occupational family, then the four closest anywhere in the index.

Read this carefully. Exposure is not displacement. A high score means current AI systems can produce this work, not that anyone will stop paying a person to do it. Adoption depends on economics, regulation and inertia that this index deliberately does not model. How the score is built.