What is happening to data scientists
- Impact
AI tools are automating data cleaning, feature engineering, model selection, and hyperparameter tuning, and generating initial insights. This shifts Data Scientists' focus towards high-level problem formulation, advanced causal inference, ethical AI governance, and strategic data storytelling.
- Risk
Significant augmentation; emphasis on complex problem formulation, model validation, and ethical data science.
The Data Scientist role will be profoundly reshaped by AI. AI will handle much of the high-volume data manipulation, routine model building (e.g., AutoML), and initial insight discovery. Data Scientists must immediately pivot to becoming experts in leveraging AI for hyper-efficiency and enhanced strategic insights, intensely validating AI outputs for accuracy and bias, and dedicating their expertise to the irreplaceable human elements of the role: profound problem formulation, nuanced causal inference, and critical ethical decision-making regarding data privacy, model fairness, and societal impact.
- Sector readiness
Leading Edge of Adoption
The data science and analytics sector is aggressively integrating AI, driven by overwhelming demand for faster model deployment, more complex insights, and scalable solutions. AI is rapidly moving beyond pilot stages to widespread adoption for data preparation, model building, and MLOps, fundamentally altering traditional workflows and skill requirements.
Where you stand
The Data Scientist role is undergoing a profound and accelerating redefinition by AI, fundamentally restructuring data preparation, model building, and insight generation.
AI will autonomously manage vast routine tasks, optimize model creation, and streamline MLOps, compelling Data Scientists to pivot to indispensable problem formulation and profound ethical governance.
Survival and impact will hinge on Data Scientists mastering AI tools, critically validating AI outputs for accuracy and bias, championing ethical AI, and providing irreplaceable human judgment and strategic insight at the heart of data-driven innovation.
What this means for you
Concrete changes to how the work gets done, in the order you are likely to meet them.
- 01
AI-Automated Data Preparation & Feature Engineering. Data Scientists are increasingly overseeing AI tools that autonomously perform data cleaning, transformation, and complex feature engineering (e.g., creating synthetic features, combining variables). This radically frees scientists from tedious data wrangling, allowing focus on novel feature ideation and model design.
- 02
AI-Powered Automated Machine Learning (AutoML). Data Scientists will orchestrate AutoML platforms that autonomously select optimal algorithms, perform hyperparameter tuning, and even generate entire machine learning models. This dramatically accelerates model building, demanding data scientists to focus on validating AI-generated models and interpreting their outputs.
- 03
Predictive Analytics for Business Outcomes. Data Scientists will leverage AI models that autonomously analyze vast and diverse datasets (e.g., sales, customer behavior, operational metrics) to predict future business outcomes (e.g., customer churn, product demand, revenue growth) with unprecedented accuracy. This enables proactive, data-driven strategy.
- 04
Generative AI for Hypothesis Generation & Insights. AI tools are assisting Data Scientists in synthesizing vast amounts of data and research, identifying hidden patterns, correlations, and even generating novel hypotheses or insights that can then be rigorously tested. This expands the intellectual landscape for discovery.
- 05
AI-Enhanced Data Visualization & Storytelling. AI can autonomously generate complex data visualizations, dashboards, and initial narratives for data science reports. This streamlines communication efforts, allowing Data Scientists to focus on refining the story, ensuring accuracy, and tailoring the message for specific stakeholders.
- 06
Focus on Strategic Problem Formulation & Framing. As AI assumes command of computational and analytical tasks, the paramount value of Data Scientists will be their irreplaceable human ability to precisely define complex, ambiguous business or scientific problems, identify the right data, and frame them for AI-driven solutions.
- 07
Ethical AI Governance & Bias Mitigation. Data Scientists will bear profound responsibility for designing and auditing AI models for algorithmic bias, ensuring data privacy, and upholding ethical standards for fairness, transparency, and accountability in AI-driven decision-making. This is a critical and paramount skill.
- 08
Human-AI Teaming for Model Development & Deployment. Data Scientists will operate in seamless human-AI teams. AI will process vast data, generate model candidates, and automate MLOps tasks. The human Data Scientist will lead conceptualization, fine-tune models, validate outputs, and manage nuanced human factors in AI deployment.
- 09
AI for Causal Inference & Explainable AI (XAI). As AI models become more complex, Data Scientists are leveraging AI tools for causal inference to understand true cause-and-effect relationships and XAI techniques to explain complex AI model predictions. This provides transparency and trustworthiness in AI-driven insights.
- 10
Continuous Data Monitoring & Model Performance Tracking. AI systems will autonomously monitor data pipelines for quality issues and continuously track the performance of deployed AI models (e.g., for data drift, concept drift, accuracy decay). Data Scientists will intervene for anomalies and manage model retraining.
- 11
Continuous Learning & Advanced AI/ML Research Literacy. The exponential pace of AI integration demands that Data Scientists commit to continuous, aggressive learning of new AI-powered tools, advanced ML algorithms (e.g., foundation models, quantum ML), and their profound capabilities and ethical implications, as a foundational competency.
- 12
Specialization in Niche AI/ML Domains. The field will see a significant rise in Data Scientists specializing in highly complex or emerging AI/ML domains, such as reinforcement learning for control, graph neural networks, federated learning, or specialized AI for specific industries (e.g., drug discovery, climate modeling).
- 13
AI-Powered MLOps Automation. Data Scientists are aggressively implementing AI-driven MLOps pipelines that autonomously manage the entire machine learning lifecycle: data versioning, model training, deployment, monitoring, and retraining. This ensures scalable and reliable AI systems.
- 14
Leadership in Data-Driven Transformation. Data Scientists in leadership roles will play a crucial role in guiding organizations through the pervasive adoption of AI, advocating for strategic data science solutions, and fundamentally reshaping the future of data-driven decision-making.
- 15
Strategic Data Storytelling & Business Impact. As AI processes data and builds models, the human skill of weaving complex analytical findings and AI-generated predictions into compelling business narratives becomes paramount. Data Scientists will focus on "prescriptive storytelling" to influence decisions and drive radical business impact.
What is pushing this change
- 01
Explosive Growth of Data (Big Data). Vast amounts of data from diverse sources (web, sensors, enterprise systems, unstructured text) provide rich input for AI models.
- 02
Advancements in AI/ML Algorithms (DL, AutoML, Reinforcement Learning). Breakthroughs in AI fields enable sophisticated analysis, autonomous model building, and intelligent predictions for complex problems.
- 03
Urgent Demand for Deeper, More Predictive Insights. Businesses and researchers need to anticipate future trends and understand causal relationships with unprecedented precision.
- 04
Increased Computational Power (GPU, Cloud). Access to powerful computing resources enables the training and deployment of complex AI models on massive datasets.
- 05
Complexity of Data Sources & Types (Unstructured, Streaming). Managing diverse, high-volume, and streaming data requires AI for efficient processing and analysis.
- 06
Need for Automated Data Preparation & MLOps. Automating data cleaning, feature engineering, and model deployment processes is critical for scaling AI solutions.
- 07
Critical Shortage of Skilled Data Scientists. The severe global shortage of experienced data scientists compels aggressive AI adoption to radically augment human capacity.
- 08
Pervasive Digital Transformation Across Industries. Companies are undergoing radical digital transformation, making robust and adaptable AI-powered solutions a central business imperative.
- 09
Ethical Scrutiny of AI Algorithms & Data Privacy. Growing concerns about algorithmic bias, fairness, and data privacy in AI models, demanding explainability and responsible AI.
- 10
Global Competition for AI Talent & Solutions. Nations and companies are investing heavily in AI to gain a technological edge, driving demand for data scientists.
Impact by sector
The headline figure is an average. Where you work changes the picture.
- Machine Learning Engineers
Highest impact; AI for AutoML, MLOps, and model deployment automation. Focus on building and managing production AI systems.
- Applied Data Scientists
AI for autonomous data analysis, predictive modeling for business outcomes, and automated report generation. Focus on applying AI to business problems.
- Research Data Scientists
AI for generating hypotheses, accelerating simulations, and exploring vast scientific datasets. Focus on fundamental AI research and theoretical breakthroughs.
- Data Architects
AI for designing data lakes, data pipelines for AI/ML, and ensuring data governance. Focus on scalable data infrastructure for AI.
- Data Ethicists
AI for auditing models for bias, ensuring data privacy, and developing ethical AI frameworks. Focus on responsible AI deployment.
Skills to build
The skills that keep the human part of this work valuable as the routine part is automated.
- 01
Statistical Modeling & Inference. Deep understanding of statistical theory, probabilistic reasoning, and applying various mathematical models to data.
- 02
AI/ML Algorithms & Frameworks. Proficiency in designing, building, and deploying AI/ML algorithms (e.g., deep learning, reinforcement learning) using frameworks like TensorFlow, PyTorch.
- 03
Problem Formulation & Business Acumen. The profound ability to identify complex, ambiguous business or scientific problems, translate them into data science challenges, and define their impact.
- 04
Data Storytelling & Communication. Expertly structuring complex analytical findings and AI-generated predictions into clear, compelling, and actionable narratives for diverse stakeholders.
- 05
Ethical AI & Explainability (XAI). Absolute mastery in identifying, mitigating, and explaining algorithmic bias, ensuring data privacy, and upholding ethical principles in AI development.
- 06
Programming & Software Engineering. Proficiency in programming languages (e.g., Python, R) and software engineering best practices for building robust and scalable AI solutions.
- 07
MLOps & Model Deployment. Expertise in deploying, monitoring, and managing machine learning models in production environments, ensuring their reliability and performance over time.
- 08
Continuous Learning & Research. A relentless commitment to continuously learning new AI techniques, adapting methodologies, and exploring cutting-edge research in AI and data science.
Tools in use
Kinds of tool worth knowing
- 01
AutoML Platforms. Platforms that autonomously select algorithms, tune hyperparameters, and even generate entire machine learning models, streamlining development.
- 02
Machine Learning Frameworks (Deep Learning). Software libraries and frameworks (e.g., TensorFlow, PyTorch, Scikit-learn) that provide the building blocks for creating advanced AI/ML models.
- 03
AI for Data Cleaning & Feature Engineering. AI-powered tools for autonomously identifying and rectifying errors, inconsistencies, and missing values in datasets, and for automated feature creation.
- 04
MLOps Platforms & Tools. Integrated platforms that automate the entire machine learning lifecycle, from data versioning to model deployment, monitoring, and retraining.
- 05
Generative AI for Data Insights & Reports. Large Language Models (LLMs) used to autonomously synthesize data, generate hypotheses, and draft initial versions of data science reports.
- 06
Causal Inference & Explainable AI (XAI) Tools. Software and frameworks that help analyze causal relationships in data and provide interpretations of complex AI model decisions.
Named tools already in use
DataRobot
VisitLeading AutoML platforms that leverage AI to automate the machine learning pipeline, accelerating model development.
TensorFlow / PyTorch
VisitProminent open-source machine learning frameworks used for building and training deep learning models.
Dataiku (AI for Data Prep) / Cleanlab (Data Quality)
VisitAI-powered data science platforms that assist with automated data preparation, feature engineering, and quality management.
MLflow / Weights & Biases
VisitOpen-source and commercial platforms for managing the machine learning lifecycle, ensuring model reliability and deployment.
ChatGPT / Claude / Google Gemini (for insights)
VisitGenerative AI models that can autonomously synthesize data, generate hypotheses, and draft data science reports.
DoWhy (Microsoft) / SHAP / LIME
VisitOpen-source libraries and frameworks for performing causal inference and for interpreting complex AI models.
In practice
Ways people in this role are already using AI, and what they get from it.
- Automate Data Cleaning & Feature EngineeringExample 1
- How
Data Scientists will implement an AI-powered data preparation tool that autonomously cleans, transforms, and generates new features from raw, messy datasets. The AI identifies missing values, outliers, and optimizes feature representations, radically reducing manual effort.
GainSignificantly reduces manual data wrangling time, accelerates data preparation, and enhances the quality and richness of features for modeling.
- Build ML Models with AutoMLExample 2
- How
Data Scientists will utilize an AutoML platform. By providing a dataset and a target variable, the AI autonomously explores various algorithms, performs hyperparameter tuning, and selects the optimal model, allowing the Data Scientist to rapidly prototype and deploy solutions.
GainDramatically speeds up model development, allows exploration of a wider range of algorithms, and enables faster deployment of AI solutions.
- Generate Hypotheses from DataExample 3
- How
Data Scientists will instruct a generative AI model to synthesize findings from a large dataset. The AI will autonomously identify hidden correlations, propose novel hypotheses, and even suggest lines of inquiry for further statistical or causal analysis.
GainSparks new lines of inquiry, expands the intellectual landscape for research, and potentially leads to groundbreaking new discoveries or business insights.
- Monitor AI Model Performance in ProductionExample 4
- How
Data Scientists will deploy an AI-powered MLOps platform that autonomously monitors the performance of machine learning models in production. The AI will detect data drift, concept drift, or accuracy decay, and automatically trigger alerts for retraining or human intervention.
GainEnsures the long-term reliability and accuracy of AI models in production, enables proactive maintenance, and minimizes performance degradation.
- Perform Causal Inference AnalysisExample 5
- How
Data Scientists will use an AI-powered causal inference tool to analyze complex observational data (e.g., from a marketing campaign). The AI will identify confounding variables and apply advanced statistical techniques to infer the true causal impact of the campaign, moving beyond mere correlation.
GainProvides robust insights into true cause-and-effect relationships, enabling more effective business strategies and policy decisions.
How this role compares
Three neighbouring roles chosen to show the direction of travel, then the roles either side of yours on the exposure scale.
- Data Analysts (Basic reporting, SQL queries) / Data Entry Clerks (Data collection)More exposed
- AI impact
Catastrophic (AI can autonomously generate routine reports; AI can autonomously capture and process data.)
Work moves toImmediate need for radical re-skilling into AI oversight, data quality management for AI, or specialization in advanced data architecture.
- AI Research Scientists (Fundamental AI) / MLOps EngineersDifferent skills, growing
- AI impact
Foundational (They design and build the core AI algorithms and systems that Data Scientists will utilize and contribute to.)
Work moves toDeep expertise in advanced AI/ML algorithms, software engineering for large-scale systems, and building robust MLOps platforms.
- Statisticians (Academic/Theoretical Focus) / Business Intelligence Analysts (Reporting Focus)Complementary, less exposed · exposure 65
- AI impact
Low-Moderate Augmentation (AI assists in theoretical exploration for statisticians; AI helps with data analysis for BI analysts), but core mathematical rigor, fundamental theorem proving, and high-level business strategy remain paramount.
Work moves toDeveloping new mathematical frameworks and theories (Statisticians); Translating business needs into reports, and dashboard design (Business Intelligence Analysts).
- 552–5 yrs
- 552–5 yrs
- 551–6 yrs
Data Scientists · this report
552–5 yrs- 601–4 yrs
- 602–5 yrs
Corporate Development Managers
602–5 yrs
Closing judgement
For Data Scientists, AI is not merely a tool but a radical force of transformation that will fundamentally redefine their role. It will autonomously handle the mundane, amplify analytical capabilities exponentially, and streamline MLOps, compelling Data Scientists to pivot to indispensable problem formulation, profound ethical governance, and strategic insight. The future Data Scientist will be a visionary orchestrator of human-AI collaboration, providing irreplaceable judgment at the heart of data-driven innovation.
Evidence and revisions
What the exposure figure rests on, what changed when it was last revised, and the published work cited for this role.
50 → 55
Window2-5 years (unchanged)
The 4 October 2026 review moved the score up by 5 points.
Microsoft's AI applicability score for the matching occupation is 0.36, in the top decile of 785 US occupations; Anthropic's observed exposure (the share of the occupation's tasks already being done with Claude) is 0.46, which is heavy by the standards of the 756 occupations measured; the US Bureau of Labor Statistics places it in the 'very high' AI-exposure tier; BLS projects employment to grow 34.6% over 2025–35. Blending our 2025 editorial figure (60%) with the 2026 evidence composite (40%) moves the score from 50 to 55.
US Bureau of Labor Statistics · Employment Projections 2025–35 and AI Exposure Categories
Official statistics · 27 August 2026AI-exposure tier: Very high. Projected employment change 2025–35: +34.6%. Matched to Data scientists.
Microsoft Research · Working with AI: Measuring the Applicability of Generative AI to Occupations
Working paper · 10 July 2025AI applicability score 0.36 (percentile 98 of 785 occupations) for SOC 15-2051.
Anthropic · Anthropic Economic Index report: Cadences
Report · 26 June 2026Observed exposure 0.46 for SOC 15-2051 (percentile 98 of 756 occupations).
UK Department for Science, Innovation and Technology · Assessment of AI capabilities and the impact on the UK labour market
Report · 28 January 2026UK context: around 70% of UK workers are in occupations with tasks AI could perform or enhance, above the US average; a one-standard-deviation rise in exposure was associated with a 3.9% fall in UK job postings.
Archived copies are served only where the licence permits; otherwise the link goes to the publisher. Full research library →
Readers' view
What people who do this work make of our reading: their own scores, their reasons, and the notes they left on each section.
Our report is one reading of the evidence. This section is the other dataset: what people who do or know this work make of it. Nobody has scored this role yet. Sign in to add yours.
—
—
55
No one has explained their score yet. A line or two about what you see in your own work is the most useful thing on this page.
No notes yet. Every section above has a “Readers' notes” line at the bottom; open one and say what you know.
Method and sources
Each report was written from a large body of published research and then, in October 2026, re-scored against occupation-level evidence: the US Bureau of Labor Statistics AI-exposure classification and 2025–35 projections, Microsoft Research’s AI applicability scores and Anthropic’s observed-exposure data, cross-checked against the reports listed in the Evidence section above. The organisations and publications below are the standing literature behind the narrative sections. Every source, with dates, licences and archived copies where we are permitted to hold them, is catalogued in the research library.
Scores are revised by blending the previous editorial figure (60%) with a composite of the three occupation-level measures (40%), capped at fifteen points per revision and rounded to the nearest five. The window shifts one notch when a score moves ten points or more. Hand adjustments are recorded with their reason in the revision log.
Research library: every source, with dates, licences and archived copies →
- World Economic Forum
- The Future of Jobs Report series — Employer survey of expected job growth and decline, skill shifts and technology adoption (2020, 2023 and 2025 editions); Artificial Intelligence and the Future of Entry-Level Work (2026).
- AI governance and transformation reports — Frameworks on ethical AI, talent strategy and industry transformation.
- McKinsey Global Institute
- AI, Automation and the Future of Work series — Research quantifying automation potential by task, sector and demographic, from "Jobs Lost, Jobs Gained" to "Agents, robots, and us" (2025).
- Industry-specific reports — Financial services, healthcare, manufacturing and others.
- PwC
- Global AI Jobs Barometer — Annual analysis of job postings and productivity by AI exposure (2024–2026 editions).
- Upskilling Hopes and Fears survey — Employee perceptions and readiness.
- Microsoft Research and Anthropic
- Working with AI (2025); Anthropic Economic Index (2025–2026) — Occupation-level usage data from Copilot and Claude conversations, the two observed-usage measures behind the 2026 revision.
- Stanford Digital Economy Lab and Stanford HAI
- Canaries in the Coal Mine? (2025–2026); AI Index Report (annual) — Payroll evidence on early-career employment in exposed occupations; annual measurement of AI capability, investment and adoption.
- Deloitte
- Human Capital Trends series — Workforce, talent and HR technology trends.
- Tech Trends series — Emerging technologies and their business implications.
- Accenture
- Technology Vision series — Forward-looking analysis of emerging technology, with emphasis on AI.
- Fjord Trends — Design, innovation and human experience in a digital world.
- Boston Consulting Group
- AI/ML insights and industry solutions — "The AI Revolution in the Workplace" and related research.
- EY
- AI and workforce reports — Adoption, talent strategy and ethics.
- IBM Institute for Business Value
- AI and automation studies — Business models, workforce evolution and leadership.
- OECD
- AI Policy Observatory — International data and policy on AI, labour markets and skills.
- Employment Outlook — Labour-market trends including technological impact (2023–2026 editions).
- International Labour Organization
- Generative AI and Jobs: A Refined Global Index of Occupational Exposure (2025) — Task-level exposure gradients for every ISCO occupation; successor to the 2023 global index.
- International Monetary Fund
- Staff Discussion Notes on AI and work (2024, 2026) — Complementarity framing: where AI augments and where it substitutes.
- UK Department for Science, Innovation and Technology
- Assessment of AI capabilities and the impact on the UK labour market (2026) — UK occupational exposure and job-posting evidence.
- Brookings Institution
- AI and automation research — Economic and social implications, displacement and skills.
- Yale Budget Lab and Goldman Sachs Research
- Tracking the Impact of AI on the Labor Market; AI and the US labour market (2026) — Aggregate labour-market monitoring; macro displacement estimates.
- Oxford University (Oxford Martin School)
- The Future of Employment — Frey & Osborne and subsequent research on susceptibility to automation.
- MIT Technology Review
- AI & Work — Reporting on AI research and its implications for industries and jobs.
- Gartner
- Hype Cycle for Artificial Intelligence — Maturity and adoption of AI technologies.
- Future of Work reports — Workplace models and talent strategy.
- U.S. Bureau of Labor Statistics
- Employment Projections 2025–35; AI Exposure Categories; Occupational Outlook Handbook — Ten-year employment projections and, from the 2025 cycle, an AI-exposure tier for every detailed occupation.
- Indeed Hiring Lab
- AI at Work Report (2025) and posting-market updates — Skill-level transformation estimates and job-posting trends by occupation.
- OpenAI
- Research papers, blog and API documentation — Large language models, generative AI, safety and societal impact.
- Google DeepMind
- Research papers and blog — Reinforcement learning, AI for science, AGI and ethics.
- Meta AI
- Research papers and blog — Large language models, computer vision, AI for social good.
- Hugging Face
- Transformers library and model hub — Open-source state-of-the-art NLP models.
- TensorFlow and PyTorch
- Documentation and community forums — Core frameworks illustrating practical capability.
- arXiv
- cs.AI, cs.LG, cs.CV, cs.CL — Pre-print research.
- NeurIPS and ICML
- Conference proceedings — Top-tier academic research.
- ACM and IEEE
- Journals and proceedings — ACM Computing Surveys; IEEE Transactions on AI.
- Kaggle
- Datasets and competition solutions — Applied machine learning on real-world problems.
- The Alan Turing Institute
- Research and reports — Responsible and applied AI.
- NIST
- AI Risk Management Framework — Voluntary framework for managing AI risk.
- European Commission
- AI Act — Risk-tiered legal framework for AI.
- Ethics Guidelines for Trustworthy AI — Principles for responsible development.
- Partnership on AI
- Research and best practice — Responsible AI development.
- AI Now Institute
- Annual reports — Social implications: power, inequality, rights.
- ACM FAccT
- Proceedings — Fairness, accountability and transparency.
- Data & Society
- Publications — Social implications of data-centric technology.
- WIPO
- Conversation on IP and AI — Intellectual-property implications of AI.
- IEEE Global Initiative on Ethics of A/IS
- Ethically Aligned Design — Recommendations for ethical AI design.
- Center for AI and Digital Policy
- Policy briefs — Accountable AI policy.