Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Atlassian research coined the AI efficiency paradox. DX's Q2 data now finds it inside engineering, with the Developer Experience Index falling

Дата публикации: 12-08-2026 13:16:50

DX has published its third quarterly report on AI's impact on engineering. This one required rebuilding the methodology from scratch. Deputy CTO Justin Reock walked me through what changed and what the numbers show about where the return on AI is going.

Основное содержимое страницы с новостью.

(DX Q2 2026 AI Impact Report)

Enterprise AI research has developed a predictable format lately with high adoption figures and rising time-savings claims galore. It can be harder to find reports that measure the friction underneath, but DX's Q2 State of AI Impact in Engineering report is one of them. DX is an engineering intelligence platform used by 500-plus organizations – Dropbox, Block, Pinterest, and BNY among them – to measure and improve how developers spend their time.

For the first time in DX's history, the aggregated Developer Experience Index (DXI) across its 500-plus customer base has moved downward. The median has slipped from 67 in Q3 2025 to 65 in Q2 2026 – a two-point drop that only looks modest until you understand where the data comes from. DX's data comes only from its own paying customers, and those customers pay for DX precisely so that DXI moves up over time. 

Justin Reock, Deputy CTO at DX and lead author of the report, discussed the findings with me over a video call:

All this data comes from aggregated anonymous metrics that are coming out of our platform, which means that they are coming from data from DX customers, which tends to bias upwards for Developer Experience Index because you have companies that are actively investing in Developer Experience by buying our platform. So our data tends to bias towards upward trend lines in Developer Experience. People buy our platform to continuously improve the Developer Experience to move their DXI driver up, and this is the first time we've ever observed this. DXI has gone down.

The two-point drop translates into hours DX has separately validated across roughly six million data points:

For every single point of improvement in the DXI, you return about 10 hours per engineer per year back to the organization, based on reduced friction, reduced toil, just less trouble in their overall development process.

Two points, twenty hours per engineer per year in new friction. That's a working half-week each engineer loses back to whatever's currently getting in their way. Across a 500-engineer org, that's 10,000 hours a year.

Reock has spent roughly seven years working on developer experience. He emphasized the DXI shift as one of the two findings in the report that caught his attention the most, because while the drop isn't dramatic, the population reporting it definitely is. 

Atlassian's Teamwork Lab – its dedicated research group on how modern teams work – surveyed 12,000-plus knowledge workers and 170-plus Fortune 1000 executives, and found in its 2026 State of Teams research that only six percent of executives could point to specific organization-wide AI ROI. The Lab called this pattern the 'AI efficiency paradox' – a situation where individual output surges, then gets backed up during reviews, approvals, and other human-judgment gates, which cancels out the flow. DX's Q2 numbers add the engineering-specific view of the same issue, with system telemetry and developer sentiment measured against each other.

Why this quarter needed a rebuild

Every previous DX report had a control group. The methodology compared cohorts of AI users against cohorts of non-users on the same metrics, identified the differences and published them. But there aren't enough non-users left to make that comparison useful, as Reock noted:

We've had to change the research game very quickly now because 95% is now our current adoption number by our reckoning when we look at API telemetry. So it's no longer really useful. There's no control group anymore.

Without a control group, DX made two changes. The first is that a dedicated data analyst joined the research for the first time, and the report authorship was paired with Brian Houck – a co-author of the SPACE research framework, who joined DX from Microsoft's Productivity team. Reock was clearly glad to have Houck on the team.

The second change is that the analysis switched to percentile-rank medians – P25, P50, P75, P90 – and structured findings against the DX Core 4 metrics (Speed, Effectiveness, Quality, Impact) and the complementary DX AI Measurement Framework. The Core 4 is DX's synthesis of DORA, the SPACE research framework, and DX's own DevEx metrics, designed to give leaders one coherent view of engineering health rather than a shopping list of individual metrics. Rather than heavy-versus-light users, the analysis now compares organizations by how far up the percentile ladder they sit on each dimension.

The adoption number also has something of a compliance problem – of that 95% telemetry figure, the remaining five percent aren't AI-free, because code they've generated using AI is still showing up in production, coming from tools outside the enterprise API telemetry stack. Reock calls it 'shadow AI' which carries an exposure risk:

That's a major PII potential risk liability GDPR liability, all these other things.

This is particularly important to organizations thinking about data-residency, sensitive-code, or regulated-industry policy.

The finding that surprised the research team

Two of the fourteen drivers inside the DXI are code maintainability (how easy the code in front of you is to understand and modify) and change confidence (how sure you feel that shipping this change won't break production). Historically the two move together. This quarter, code maintainability rose 3.8% while change confidence dropped 6.1%.

Reock explained that when Gratiana Fu, DX's research analyst on the report, first surfaced the chart, his response was to send it back:

The first thing I said was, 'run it again'. We've never seen that tension before. This is novel. We need to make sure that there's not something wrong in the way that we've done this calculation. And then of course there wasn't.

When the numbers came back the same, the team got excited. Reock said they'd been out at roughly 20 conferences over the previous quarter and by his read the industry keeps saying the same (correct) things about deriving value from AI and taking care of quality. The divergence of maintainability versus trust was the first thing they'd found that felt new. He described the feeling as being like an earthquake, and explained what the finding means:

The code is more maintainable. Engineers are able to approach the code because they have AI assistants, agents, whatever that is, but they trust it less.

The volume of code involved adds weight to this. Median Pull Request (PR) size has doubled year-on-year – from 42 lines in Q3 2025 to 72 in Q2 2026. Bigger pull requests are harder to reason about, harder to review, and harder to safely revert. Reock has been writing code professionally since the late 1990s, and he has a strong view on where this trend comes from:

It was a mark of honor to be able to implement a use case with as little code as possible. It's called an elegant solution, and AI produces average, mediocre code, not elegant code.

He was careful to say that the mediocre-code story isn't the only explanation. There's a build-pipeline behavior that predates AI and is now getting amplified by it:

What engineer is going to submit five PRs and wait five hours for five features when they could just ship one PR with five features in it, and then let the single build pipeline complete?

Larger pull requests are partly a code-generation story and partly a symptom of build times that were already too long – AI made batching easier and lowered the cost of doing it.

What engineers and their managers are actually experiencing

Taken together, the DXI's individual driver movements describe what daily engineering work now feels like. Documentation went up sharply – Reock called it the report's one clear bright spot, because AI is good at generating and maintaining docs and most engineers dislike writing them from scratch. Code maintainability and production debugging improved slightly, but everything else moved down.

Deep work is down because engineers are context-switching between multiple agents – Cursor, Copilot, Claude Code, the various assistants on the desktop. Build and test perception is down because everything upstream is now faster, so any part of the pipeline that hasn't been optimized feels disproportionately slow by comparison. Cross-team collaboration is down too – Reock explained:

Nobody's talking to engineers anymore. They're asking agents for problems, and it just keeps the conversation very isolated.

Incremental delivery took the biggest hit of all the drivers, which matches the PR-size doubling. Engineers are shipping bigger, less-frequent, less-testable batches of change, which then create longer review cycles, more complicated merges, and slower recovery when something breaks.

This is also supported by the data around change failure rates. According to DevOps Research and Assessment (DORA), the benchmark for the percentage of feature deployments that require immediate remediation is four percent. This quarter's data shows some organizations moving three percentage points above their previous baseline, which translates into 75% more defects for those teams. That's not the median experience – the distribution peak still sits near zero, so most organizations haven't seen the metric move quarter-over-quarter. The range now extends past ±3 percentage points, where the previous quarter's sample stayed inside ±2. Most organizations haven't moved. The best teams and the worst teams have pulled further apart. Reock believes that AI is amplifying whatever quality practices were already there, noting:

This is nice data to back up this statement that everyone's been saying: 'Oh, AI amplifies the good and the bad.' Well, here it is. If you already had bad quality process, it's getting even worse.

Reock has seen straw polls of engineers asked to describe how they feel about AI in three words. The pair that came up most often, in the same answer: 'excited' and 'terrified'. That kind of daily emotional state is being worked through by a lot of engineers and their managers right now, when it has already been a hard year - which is a more important way to think about the decline of the DXI as much bigger than a two-point median. There are real people behind it undergoing a lot of friction. 

On the positive side, engineering managers are now writing four times more code than they were a year ago. I wasn't sure if this was a plus, but Reock interprets this finding as a return of the player-coach model DevOps used to celebrate:

Engineers have been complaining for years that we have hour-long build times and nothing's getting done, and all of a sudden engineering managers are trying to submit PRs and they're like, 'We have hour-long build times. This is terrible. Something should be done.' So something's getting done, hopefully.

Product managers and designers are now showing up as authors of code merged into production. Reock tells me this comes from the same code assistants engineers use, Cursor, Copilot, and Claude Code among them, rather than Lovable or Figma MCP. Whatever it eventually means for how engineering organizations think about role definitions and career paths, PMs and designers are now producing code that has to be reviewed, tested, and maintained by someone else.

Calculating costs

Median quarterly organizational AI spend has gone from roughly $1.5K in Q3 2025 to nearly $44K in Q2 2026 – a 28-times increase, calculated across input, cached-input, and output tokens at frontier-model rates. Every industry vertical shows the same shape.

The one measurement that hasn't kept pace is innovation ratio, where the share of engineering time spent on new capabilities rather than maintenance and operational work. It moved from 57% in Q3 2025 to 58% in Q2 2026. 

Reock emphasized caution about drawing hard conclusions from a single quarter's innovation ratio number. Some of the organizations in the sample may be actively choosing to spend their AI-freed time on tech debt, testing, and quality work rather than net-new features – which would suppress this metric in the short term. He was clear, though, on why:

We clear the bottleneck. We exploit the code generation bottleneck. We exploit the PR bottleneck, but we're just seeing that then create new constraints, new bottlenecks, and effectively these wins are just absorbed by other sources of friction and toil in the environment.

The time savings data supports this – median self-reported time saved from AI assistants has climbed from 3.3 hours per week in Q3 2025 to 6.1 hours per week in Q2 2026, with daily heavy users tracking higher still. Those saved hours are not showing up at the portfolio level as new-feature output. In the DX report chart workflow-time categories are ranked by annualized cost across the 400-plus company sample. Meeting-heavy days come first, interruption frequency second, and only then does the AI time savings offset appear – ahead of build-and-test wait time, dev environment toil, review wait time, code comprehension, and information-seeking. AI is saving time, but it's a smaller amount, and there are five other categories of friction it hasn't touched at all.

This is where DX's engineering data reflects the pattern documented by Atlassian's Teamwork Lab. DX's contribution is a granular view, with system telemetry and developer sentiment at a resolution fine enough for a leader to reason about their own environment rather than the wider workforce average.

Where the research is going next

Reock also discussed a metric DX is developing which will be interesting to see in future, called 'AI effectiveness'. The mechanism is a local daemon – a piece of software running in the background on the developer's machine – that observes full developer-agent sessions, then plays those sessions back to the agent for self-assessment against the organization's context and the scope of the work. He elaborated on why it's important to measure:

When we can trace a change failure to production where a certain threshold percentage of it was generated by AI, what was the use case? How was the agent experience? Because then we can try to work out if we can link lower quality to lower agent effectiveness, poor context, bad scope, bad steering, and to specific use cases that are highly problematic.

This is still in controlled release. Session-level observability of coding agents is now well-established – Langfuse and Coralogix trace the major assistants including Claude Code and Cursor. What DX is proposing is applying that observability inside a developer experience framework, correlating conversation quality with production defects. DX is also moving toward what Uber calls 'feature velocity' – a mix of PR throughput and innovation ratio designed to measure value shipped rather than motion.

Recommendations for leaders

Reock's advice for leaders comes down to four things, starting with managing expectations. The industry is not seeing 2x, 5x, or 10x productivity gains, but single-digit percentage improvements per quarter with room to compound. A DX longitudinal study from February found median PR throughput up 7.7% and mean throughput up 13% over 16 months, and Reock recalled:

I remember a time when we would have celebrated 8 to 13% gains. Now with all the hype around AI, it's like no, 2x, 5x, 10x.

He carries the same view into how he talks to leaders about their own numbers:

If you're coming into this expecting a 2x output, a 5x output, a 10x output, and you're not seeing it, it's not because you're doing anything wrong. It's not because your teams are doing anything wrong.

The second piece of advice is to map value streams, now that code generation is no longer a binding constraint. Third, invest in learning time, and Reock is keen to note that this should not be in the form of manuals or wikis, but actual dedicated hours for engineers to develop the skill of working well with agents. Finally, protect psychological safety and be explicit about why AI is being brought into the organization. His anchor example for how to do this well is Zapier, and he lit up when he talked about them:

They're hiring more than they have in the history of their whole company. I think that's the biggest metric for me. They're saying we can get more value out of any single engineer, more ROI out of every hire. Why wouldn't we be hiring like crazy right now?

Cutting headcount because AI increased capacity is a huge mistake. If anything, Reock believes the data supports increasing capacity per engineer, then hiring more of them because each one is now more valuable:

The voice of the engineer, I think, has just never been more important.

My take

Engineers can read AI-generated code but they don't trust it. Leaders running an agentic rollout have to explain why code is harder to maintain and confidence in change is going down in their own environment – and what they're doing about it. "More evals" is a starting point. "We haven't looked" is not.

DX's data shows non-AI friction still costs more time than AI saves. New engineers get productive faster, but there's still a lot of friction that is taking a toll on people, including context switching, isolation from colleagues, and larger changes to review. This is what happens when a whole industry adopts AI faster than the processes around it can adjust. For everyone's sakes, good leaders need to ask what is making work harder and actually do something about it.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1the diginomica network - an inside view of how one CIO rebuilt software development around AI and people07.1425-09-2026
2AI coding adoption rate hits 97%, Black Duck study reveals012.1709-06-2026
3Google-Analyse zeigt Verdopplung von Softwareschwachstellen durch KI012.6630-09-2026
4AI Made Engineering Faster. Why Not The Business?06.6228-09-2026
5Executive Intelligence podcast - Workday's Gerrit Kazmaier on facing down the SaaSpocalypse, and why Workday had to change before agentic AI took hold05.1530-09-2026
6the diginomica network - an inside view of a CIO replacing tools with in-house AI builds017.825-09-2026
7Study examines risks companies face when relying too heavily on AI systems08.8228-09-2026
8Why AI won't fix Britain's supply chains07.3524-09-2026
9Atlassian soars as the ‘SaaSpocalypse’ and tokenomics crises fail to prevent a strong year-end with context as king011.0209-08-2026
10AI Coding Assistants in 2026: Avoiding Pitfalls and Maximizing Value010.522-05-2026

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 7.35. Источник: diginomica.com.