AI Observability Uncomfortable Report Card: $74 Million Outages and Agents Running Blind

AI Observability Uncomfortable Report Card: $74 Million Outages and Agents Running Blind

New Relic has released its 2026 Observability Forecast, and the headline number is a familiar kind of shock. The businesses surveyed lose an annualized $74 million to high-impact IT outages — $1.85 million per hour, $30,833 per minute. The figure is barely down from last year’s $76 million.

Its base is large: 2,575 engineering and IT leaders across 24 countries and 12 industries, surveyed with Enterprise Technology Research in April and May. Outages remain routine. Thirty-six percent of respondents report a high-impact outage weekly or more. Detection takes 41 minutes on average; resolution another 54.

The paradox sits at the center of the data. AI adoption is the top driver of observability spend for the second year running — and a quarter of organizations run AI agents in production with no monitoring at all. The technology forcing everyone to buy monitoring is itself the industry’s biggest blind spot.

New Relic’s coinage for the wasted motion is “phantom velocity”: teams deploy fast, then lose the gains to emergency fixes. Engineers now spend 37 percent of their time on disruptions, up from 33 percent. Forty-two percent of organizations still learn about outages through manual checks or customer complaints.

There is also a vendor story under the numbers. This report comes from a company selling the cure it describes, amid a fierce market race. Reading it that way is not cynicism. It is due diligence.

The Competitive Picture for AI Observability

To judge them, look at who publishes this report and why AI observability became the market’s fastest-growing corner.

A market growing around its own disruption

The observability market sits near $3.35 billion in 2026, on projections toward $6.93 billion by 2031. The AI observability sub-segment grows faster still — analysts peg it near 25 percent annually through the decade’s end. Datadog says its LLM observability customer count more than doubled in six months, and its AI-native cohort now represents a meaningful share of revenue.

Consolidation is already underway. Snowflake acquired Observe in January; ClickHouse bought LLM-observability leader Langfuse the same month; Mintlify absorbed Helicone in March. When a sub-segment consolidates that fast, the land grab has ended.

New Relic’s position is the subtext

New Relic is private — Francisco Partners and TPG took it off the NYSE in November 2023 at roughly $6.5 billion, and it has published no financials since. Its last public year showed $925.6 million in revenue, growing 18 percent, while rival Datadog’s 2025 revenue reached $3.3 billion.

Independent spending data is unkind. Enterprise Technology Research’s own observability survey — the same firm New Relic partnered with — scores New Relic’s net spending momentum at just 3 percent, lowest among tracked vendors. Microsoft, Datadog, and Dynatrace lead that measure. The Observability Forecast is, among other things, a demand-generation instrument for a company defending position.

The product strategy is real regardless. New Relic shipped an SRE Agent in February, AI Agent Monitoring the same month, and published an internal-data AI Impact Report in January. The company is betting heavily on agentic operations.

OTel commoditizes the floor

OpenTelemetry keeps eating the collection layer. Nearly three-quarters of organizations are standardized on it, migrating, or testing — only 2 percent have ruled it out. The AI sub-segment is standardizing on OTel’s gen_ai semantic conventions too, which means telemetry capture is becoming a commodity.

That pushes vendors up the value stack: from collecting data to reasoning about it. AI observability is where that reasoning now differentiates — which is precisely why every vendor publishes reports like this one.

What the Data Shows

The report’s internal story is one of expensive stasis. Outage costs barely moved year over year. Outage frequency did not improve — 36 percent still get hit weekly. Engineers spend more time firefighting than last year. AIOps has reached only 36 percent of organizations, though another 35 percent plan to adopt within a year.

The AI-era findings carry the new weight. Eighty-three percent agree AI-generated code makes observability essential, yet fewer than half deploy AI application observability. One in four runs agents blind. The report’s own logic: an agent that changes configurations without monitoring is “a production incident waiting to be discovered by a customer.”

The ROI findings deserve caution. Forty-two percent of organizations monitoring agents report a 3x-plus return on observability — double the 21 percent reported by agent-deployers that do not monitor. But correlation is not causation. Organizations mature enough to monitor their agents are likely better run in general. The survey is also vendor-commissioned, and last year’s $76 million figure was a median while this year’s $74 million is a mean — a quiet shift that makes the year-over-year comparison imperfect.

What’s New vs. Repackaged in This AI Observability Forecast

Genuinely new

The agent blind-spot finding. One in four production agents running unmonitored is a genuinely new statistic. The monitor-versus-don’t ROI split is the report’s sharpest analytical cut. “Phantom velocity” is also a fresh, useful frame for the deploy-fast-then-fix cycle.

Improved

The ROI case is more granular than last year — 73 percent of agent-monitors report at least a 2x return, and 70 percent of all adopters report faster detection since adopting observability. Meanwhile, OTel adoption data now tracks the standard’s move from experiment to default.

Repackaged

The $70-million-plus outage figure is the same shock statistic the report has led with for years. It moves a few million either way annually. And the “AI is the top driver of spend” headline is a repeat of last year’s finding. The firefighting-time figure has moved from 33 to 37 percent — worse, not new.

Unclear

How one quarter of agents run unmonitored while AI monitoring usage keeps growing — the report does not reconcile the two. What “3x return” is measured against. And whether the median-to-mean switch in the cost figures was methodological necessity or framing.

The Question That Wasn’t Answered

If observability spend keeps rising — 72 percent of organizations expect increases — and outage costs stay flat, what exactly is the money buying? The report celebrates returns at the organization level while its own topline shows the market-level bill has not moved. That gap is the question.

A recursive one follows it. Observability vendors now ship their own agents — New Relic’s SRE Agent, Datadog’s Bits AI — that triage incidents autonomously. Who monitors those? The industry selling “never run an agent blind” has begun running agents of its own.

What This AI Observability Report Means for You

If you lead engineering, treat the 25 percent figure as your checklist: instrument your agents before you scale them, not after. Standardize on OpenTelemetry so your telemetry survives vendor switches. And when a vendor cites survey ROI, ask for cohort data you can verify — the monitor-thy-agents return is plausible, but it rewards maturity as much as tooling.

If you evaluate observability platforms, read this report as a market map, not a scoreboard. The consolidation wave — Snowflake-Observe, ClickHouse-Langfuse — signals where the category is heading: AI monitoring folded into data platforms. Shortlist accordingly, and weigh exit risk when your monitoring vendor itself gets acquired.

If you compete in this market, note what the report telegraphs. Every vendor now argues agents need watching; the differentiator has moved from telemetry to autonomous reasoning. The next battleground is who monitors — and remediates — first.

AI Observability Uncomfortable Report Card: $74 Million Outages and Agents Running Blind

Editor’s Note

This article draws on New Relic‘s 2026 Observability Forecast and its September 2026 release, prior-year report materials, and independent market research from ETR, Mordor Intelligence, and industry analysts. All survey figures are self-reported by respondents and commissioned by a vendor that sells observability products; outage cost estimates are not audited financial data. No product was evaluated hands-on by TechRecast.