Only 6% of respondents qualify as AI high performers, defined by McKinsey as those who attribute at least 5% of EBIT to AI and describe its impact as significant, according to McKinsey’s 2026 research, The State of AI in 2026: On the Road to ROI. A separate HBR Analytic Services survey, sponsored by Cloudera, Taming the Complexity of AI Data Readiness, puts the share of enterprises with data they consider fully ready for AI at 7%.

Taken together, the reports illustrate a broader enterprise AI gap between experimentation and accountable production. Many enterprise AI initiatives live somewhere between agentic experimentation, where models operate on a short leash and run in a controlled setting, and production, where agents make or influence real decisions under real-world stress and compliance and audit scrutiny. The chasm is the leap from experimentation to accountability.

Moving from experimentation to production means the business must accept consequences: customer-facing outputs, financial decisions, regulatory filings, or clinical recommendations the organization must defend. When Air Canada’s chatbot misquoted its bereavement-fare policy to a customer, a tribunal held the airline liable for negligent misrepresentation and rejected the argument that the chatbot was a separate legal entity responsible for its own statements. The case illustrates how accountability can attach once an AI system begins speaking or acting on an organization’s behalf.

The cost of getting agentic AI wrong is no longer the cost of a failed pilot; it could be reputational damage, regulatory action, or bad decisions executed at scale.

Shipping From Pilot to Production

So how does an organization know when an agent is ready to carry that weight? Practitioners who move enterprises across the line tend to converge on the same test, and it has less to do with the model than with the data beneath it. Before any agent touches a customer, a ledger, or a filing, they say, the organization must be able to answer basic questions about the fields it will read and write: where they come from, who owns them, how clean they are, and who is accountable when they’re wrong. Most can’t.

Jeremy Carmona, founder of Clear Concise Consulting, describes the floor from an operational perspective based on his assessments across nonprofit, healthcare, and government clients: definitions for the fields the AI will touch, a known and low duplicate rate, ensuring required fields are actually populated, and assigning a person to own the data who can refuse data requests. “Almost nobody clears it cold,” he says. The organizations that do were typically burned by an ugly migration years earlier and never wanted to feel that way again.

Amine Badaoui, senior technical product manager at Rackspace Technology, provides a five-question diagnostic his team uses before recommending an AI deployment for production:

  • What are the authoritative data sources?
  • Who owns them?
  • How is quality measured?
  • Who has access?
  • How do changes get tracked?

If those questions create uncertainty or conflicting answers within the organization, foundational work remains incomplete before moving forward.

Yasmeen Ahmad, formerly managing director of Data Cloud at Google Cloud and now senior vice president and general manager of Workday Data Cloud, specifies a four-component framework for production AI readiness: a universal context engine, a shared, governed layer that connects enterprise data to its business meaning, relationships, permissions, and provenance, allowing AI agents to interpret terms such as “customer,” “exposure,” or “revenue” consistently across systems while respecting the policies governing their use; a unified cross-cloud data estate where agents can reason securely across environments without hitting incompatible systems; real-time operational access, because batch ETL introduces latency that disqualifies time-sensitive use cases; and embedded explainability and audit infrastructure, because governance can’t be bolted on after the fact.

If a compliance team can’t trace what data an agent touched, what logic it applied, and what transaction it executed, Ahmad says, that agent will not be allowed to operate autonomously at scale. Most organizations, she notes, are missing at least two of the four readiness prongs she cited.

Dost Mushtaq, CEO of PreMetric, which works at the pre-deployment stage helping organizations decide whether AI initiatives should proceed, be modified, or stop before becoming embedded in the business, adds that getting the environment ready for deployment doesn’t mean it has to be enterprise-ready. “What kills AI programs isn’t starting small. It’s starting without a reliable foundation and discovering the problem after the commitments are made,” Mushtaq says.

What should that foundation look like? Bob Bloem, senior director and head of data and analytics at cloud-services provider Caylent, describes that floor this way: maintained data definitions for critical domains, documented lineage at least to the field level [that is, a record of every individual data column’s journey from source system through each transformation to wherever it’s consumed] for data feeding the AI system, active quality monitoring with agreed service level agreements, a functioning data catalog with actual adoption rather than just installation, and clear data ownership assigned to humans who are measured on it. That is the floor, he notes: it does not (yet) include more mature capabilities like automated data contracts or real-time quality scoring. In his experience, fewer than one in ten organizations meet it before launching an AI initiative, consistent with Gartner’s finding that 63% of organizations either don’t have or aren’t sure they have the right data management practices for AI.

Sylvie Ouziel, CEO of Blue Bridge Group AI, argues the blanket readiness prerequisite is being overapplied. For specific GenAI use cases such as contract abstraction, regulatory compliance checking, and document-based automation, the system operates on existing documents rather than training data drawn from enterprise history. In these cases, the data hygiene readiness bar is use-case dependent, and organizations applying the same standard to document-reasoning applications as to ML models trained on enterprise data are overshooting.

Practitioners who navigate this most effectively treat it as triage: identify the specific data domains feeding the first production initiative, get those data domains production-ready, prove the model, and expand from there.