How Agentic AI Is Transforming Work in 2026

How Agentic AI Is Transforming Work in 2026: The Evidence From OpenAI’s Codex Research

The debate about whether AI will change how people work is over. The question now is how fast, how completely, and what organizations need to do to capture the benefits rather than absorb the disruption.

In June 2026, OpenAI published what is likely the most detailed empirical study of agentic AI adoption at work yet produced: a research paper titled “The Shift to Agentic AI: Evidence from Codex,” authored by Drew Johnston and David Holtz, analyzing usage data across three distinct populations — external personal-account users, external organizational users, and workers within OpenAI itself. The paper was accompanied by a separate OpenAI report on research acceleration published in September 2026, which updated the internal data through mid-August and revealed patterns that extended and reinforced the earlier findings.

Together, these two documents present the clearest empirical picture available of what happens to work when agentic AI tools become the primary interface between people and AI systems. The picture is not subtle. It describes a transformation that, at OpenAI internally, is already essentially complete — and that is now propagating outward to the broader enterprise market at accelerating speed.

The Core Finding: Agentic AI Changes What Gets Delegated, Not Just How Fast

The central finding of the Codex study is one that distinguishes agentic AI from all prior generations of AI tools. Earlier AI tools — including conversational AI like the original ChatGPT — primarily made existing tasks faster. People did the same things, but with AI assistance that reduced the time required per task.

Agentic AI changes what gets delegated entirely. According to the Codex study, as users became more experienced with Codex, they progressively moved from asking it to complete simple, short tasks to delegating complex, multi-step, long-horizon work. By May 2026, 80.6 percent of sampled individual Codex users had made at least one request estimated to require more than 30 minutes of human work. 70.2 percent had made at least one request estimated to require more than one hour of human work. And 25.6 percent had made at least one request estimated to require more than eight hours of human work.

This is not the profile of a tool being used to speed up existing tasks. An eight-hour human task is not a faster version of a two-minute task — it is a qualitatively different category of work. Users delegating eight-hour tasks to Codex are not accelerating their workflows; they are outsourcing entire work sessions to an agent and receiving the output. This represents a fundamentally different relationship between human worker and AI tool than anything that existed before agentic AI systems were available.

The research team connected this to a broader finding about workflow structure. Studies by Demirer and colleagues, referenced in the Codex paper, found that AI generates the largest productivity gains when it can execute contiguous chains of activities. The importance of workflow structure and task adjacency explains why agentic AI outperforms conversational AI by such large margins on complex work: agents can maintain context and execute across long chains of connected tasks, while conversational AI requires human reengagement at each step.

The Numbers: Five Times the Users in Six Months

The scale of Codex adoption in the first half of 2026 was exceptional by any measure. The number of active Codex users grew more than fivefold between January 1 and June 1, 2026 — in six months. This growth rate was broad-based, but the study documented that the most rapid increase was occurring outside Codex’s initial core audience of software developers. Non-engineering functions were adopting agentic AI tooling faster than engineering was growing.

The pattern of adoption within OpenAI itself provides a preview of what this adoption trajectory looks like as it completes. Engineers at OpenAI began adopting Codex first, gradually, in late 2025. The average OpenAI engineer shifted the majority of their AI tool usage to Codex by December 2025. By June 2026, the average engineer was generating 99 percent of their output tokens through Codex rather than through ChatGPT. Conversational AI had not disappeared — it accounted for the remaining 1 percent — but it had been essentially replaced for substantive work.

The non-engineering functions followed a similar trajectory on a delayed timeline. Legal, finance, and recruiting crossed the majority-Codex threshold around April 2026. By June 2026, as of the paper’s data collection cutoff, Codex had become the primary AI tool for every department at OpenAI. The full transition from conversational AI to agentic AI as the primary human-AI work interface was essentially complete across the entire organization within a period of less than a year.

The output implications of this transition were dramatic. In June 2026, the median OpenAI employee in a legal role was generating 13 times more monthly output tokens across Codex and ChatGPT combined than they had in November 2025 — seven months earlier. The median researcher was generating more than 50 times as many output tokens. Across all functions tracked, the median worker’s output token count rose at least ten-fold between November 2025 and June 2026.

Research Acceleration: 3.1 Agent-Workdays Per Human Researcher

OpenAI’s September 2026 research acceleration report, published alongside GPT-6 Astra’s launch week, extended the Codex study’s findings specifically for the research organization through mid-August 2026. The headline metric is one of the most significant data points about agentic AI’s impact on knowledge-intensive work yet published: as of mid-August 2026, the OpenAI research organization was using 3.1 agent-workdays of effort for every workday of human labor.

Before June 2026, total agent runtime across the research organization was still below that of total human labor hours. That relationship had inverted by June, and continued to widen through August. By mid-August, agents were doing more than three times as much work, measured in workday-equivalent effort, as the human researchers they were supporting.

The cost dimension of this acceleration is equally striking. The median researcher in OpenAI’s research organization was using more than $600 per day of AI inference at API prices by mid-August 2026. At the 90th percentile, individual researchers were consuming more than $7,000 in inference tokens per day. These figures reflect the level of agent usage that the research organization found necessary to maintain competitive research output at the frontier — a data point for enterprise organizations thinking about what “serious agentic AI deployment” actually costs per user when fully adopted.

The experiment count data reinforced the output picture. The number of experiments per active experimenter hit an all-time high in August 2026 — the highest since OpenAI began tracking this metric in January 2025. The increase correlated with Codex adoption, though the company noted that available compute had also grown, making it difficult to fully isolate the Codex effect from the compute effect. Multiple teams that had previously held office hours to help researchers troubleshoot experiments reported declining attendance in 2026; one team stopped holding office hours entirely to focus on other system improvements. The troubleshooting function that human-to-human office hours had served was being handled by agents.

Task Complexity: The Shift From Minutes to Hours to Days

One of the most important dimensions of the Codex study is what it reveals about the direction of task complexity over time. The headline user growth numbers could, in principle, reflect a large population of new users making simple, short requests. The complexity data shows the opposite.

Since the beginning of 2026, the share of individual Codex users who submitted at least one request for a task estimated to require more than eight hours for an experienced human to complete increased nearly tenfold. This is not a marginal shift at the edges of the user distribution. It reflects a broad and accelerating movement toward delegation of very long, complex tasks across the user population.

Among daily active users at OpenAI, the 99th percentile users by mid-2026 were regularly generating more than 60 hours of Codex agent turns per day, distributed across multiple parallel agents. A single human worker directing agents that collectively run 60 hours of work per day is operating at a leverage ratio that has no historical precedent in knowledge work. One person, effectively, directing a team of agents that could accomplish more than a week of human-equivalent effort in a single workday.

More than 10 percent of Codex users were managing three or more concurrent agents at some point each week. And 26.6 percent had used “skills” — structured instructions for complex workflows that users create and share — reflecting the emergence of a layer of workflow codification and knowledge sharing built on top of agentic AI. Users were not just using agents; they were building reusable agent architectures and sharing them with colleagues.

Beyond Engineering: Legal, Finance, and Research Transformation

The popular narrative about AI agents focuses heavily on software engineering. The Codex data tells a more complete story where engineering is the early adopter, not the exclusive beneficiary.

Research saw the largest output multiplier in the data: the median OpenAI researcher’s monthly output token generation was 56 times higher in June 2026 than in November 2025. Customer support rose 32 times over the same period. Engineering rose 27 times — substantial, but behind both research and customer support in relative terms. Legal grew more gradually but still reached 13 times the November 2025 level by June 2026.

The fact that research showed the largest multiplier is significant. Research is the domain where human judgment, creativity, and domain expertise are most clearly essential — and yet it is also the domain where agentic AI produced the largest output acceleration at OpenAI. This is likely because research involves exactly the kind of long, contiguous chains of connected tasks — designing experiments, running them, analyzing results, generating follow-up hypotheses, testing them — where agentic AI’s ability to execute across extended workflows provides the most leverage over a conversational tool that requires human reengagement at each step.

The quantum computing experiment workflow documented in the September 2026 reporting illustrated this architecture concretely. A researcher used GPT-5.6 Sol with Codex to autonomously run quantum computing experiments, analyze results, and calibrate qubits — a closed-loop experimental workflow where the agent orchestrated the full cycle from experiment design through instrument interaction to result analysis, with no human in the execution loop. The human contribution was in setting up the architecture, defining the objective, and reviewing final outputs. The agent handled every intermediate step.

The Implications for Workforce Organization

The Codex study’s conclusion section addressed what the researchers described as the implications for productivity, job reorganization, and workforce restructuring — three distinct phenomena that the data illuminates in different ways.

On productivity, the output data is clear: agentic AI tools produce large, measurable increases in output per worker when adoption is deep. The 10x to 56x increases in output token generation across different functions at OpenAI are not measurement artifacts — they reflect real work being done that would not have been completed without agents.

On job reorganization, the study’s findings suggest that the primary change is in what humans do, not whether humans are involved. The tasks that agents handle independently — executing code, researching specific questions, drafting initial versions, running experiments — were previously done by humans alongside other work. When agents take those tasks, human attention shifts toward higher-level direction, quality review, and the design of new agent workflows. This is reorganization, not elimination — at least in the current phase of adoption.

On workforce restructuring, the data is more ambiguous. The study documented that as Codex usage increased within OpenAI, the functions of certain human roles changed substantially. The office hours for experiment troubleshooting that stopped being necessary is one example. Whether the human capacity freed by agentic AI tools is redeployed to higher-value work or results in workforce reduction depends on organizational decisions that the technology itself does not determine.

McKinsey’s August 2026 survey corroborated the Codex study’s findings at a broader scale. Thirty-nine percent of organizations reported expecting AI-driven workforce changes in 2026, up from 32 percent in 2025. The direction of change — toward higher-level human work and reduced involvement in task execution — was consistent across both datasets.

What Enterprise Organizations Should Take From the Codex Data

For enterprise leaders evaluating or expanding agentic AI deployments, the Codex study and research acceleration report offer several specific, evidence-based points of reference that are more useful than the general claims that characterize most agentic AI discourse.

The adoption trajectory at OpenAI — engineering first, other functions following six to nine months later, full organizational transition within a year — provides a rough template for what enterprise agentic AI adoption might look like in organizations that commit to it seriously. The timeline will vary, but the sequence and the completeness of the transition in OpenAI’s case suggests that organizations currently in early phases of agentic AI adoption should plan for a more complete organizational transformation than a typical software rollout requires.

The cost data — median researcher at $600 per day, 90th percentile at $7,000 per day in inference costs — provides a realistic benchmark for budgeting serious agentic AI use at the individual level. These figures reflect frontier research work at one of the world’s most AI-intensive organizations, and typical enterprise use cases will be less intensive. But the order of magnitude is relevant for finance teams that are still modeling AI costs based on per-seat SaaS pricing rather than usage-based inference costs.

The complexity shift data — the tenfold increase in users submitting eight-hour tasks — suggests that measuring agentic AI adoption by user count understates the actual change in AI utilization. Organizations that count active users as the primary adoption metric should supplement it with task complexity and output volume metrics to capture the full picture of how agent usage is evolving.

And the skills data — 26.6 percent of users building and sharing structured workflow instructions — indicates that the most effective agentic AI users are not just using agents but building reusable workflow assets. Organizations that invest in enabling and incentivizing this kind of workflow development are likely to compound their agentic AI returns faster than those that treat agent use as a purely individual activity.

Conclusion: The Transition Is Underway and Accelerating

The OpenAI Codex study and research acceleration data together provide the strongest empirical evidence yet that the transition from conversational AI to agentic AI as the primary form of human-AI work collaboration is not a future event. It is happening now, at accelerating speed, and the organizations where it is furthest along are showing productivity multipliers that would have seemed implausible two years ago.

Within OpenAI, the transition is essentially complete. Across the broader enterprise market, it is at an early stage — but the 5x user growth in six months, the expansion beyond engineering into legal, research, and customer support, and the rapid increase in task complexity all indicate that the trajectory established at OpenAI is not unique to one organization. It is the direction of travel for knowledge work more broadly.

For organizations that are still treating agentic AI as an experiment to be evaluated, the Codex data suggests a different posture: this is a transition that is already underway, and the question is not whether to engage with it but how quickly and how deliberately. The organizations that build the workflow infrastructure, the governance architecture, and the reusable skills libraries now will compound those investments as the capability of the underlying models continues to improve. Those that wait for a clearer picture may find that the transition has already happened around them.

Follow our site for the latest research and analysis on agentic AI, the future of work, and enterprise AI strategy.

Leave a Comment