The world's conversation about artificial intelligence centers almost entirely on models. Which foundation model is largest, which benchmarks it has broken, which lab has announced the next capability leap. What that conversation consistently obscures is a more fundamental truth: an AI system is only as intelligent as the data available to it. Models are the engine. Data is the fuel. And for enterprises whose most valuable data cannot safely leave their own walls, a cloud-based AI strategy is, by definition, a constrained one.
Daniel Glogowski, principal product manager at NVIDIA, has heard this realization arrive in boardrooms with increasing urgency. "What we're seeing is that enterprises have genuinely sophisticated AI ambitions," he says. "The models exist. The use cases are clear. What's blocking them isn't capability, it’s the data. The most valuable data is the data they're least able to activate."
That tension is becoming impossible to ignore. Legal firms, for example, are now confronting a stark reality: client information shared with cloud-based AI services is potentially discoverable in litigation. Entire categories of AI-assisted work, including contract analysis, case research, and due diligence, become legally untenable the moment sensitive client communications or privileged documents touch an external AI platform.
The same structural problem applies across financial services, healthcare, insurance, and any enterprise that holds proprietary or regulated data. The most valuable AI use cases are precisely the ones that require the most sensitive data, and the most sensitive data is exactly what cannot go to the cloud.
A new survey of 203 enterprise IT decision-makers, conducted by Cloudian in February 2026, confirms the scale of this reckoning. Nearly 79 percent of organizations have already moved some AI workloads from public cloud to on-premises or private infrastructure, or are actively doing so. A further 73 percent plan to shift more workloads on-premises or expand hybrid deployments over the next 24 months. Three forces are driving this: the data sovereignty imperative, the unpredictable economics of cloud AI, and the performance demands of production workloads that cannot tolerate hyperscaler latency.
The data problem that policy alone cannot solve
Of the three drivers, data sovereignty has moved fastest from a compliance consideration to a boardroom priority. The survey found that 74 percent of respondents characterize shadow AI, meaning the unauthorized use of cloud-based AI tools by employees, as either a critical or significant data security concern. Nearly one in four organizations has already experienced documented incidents of employees uploading confidential data to external AI services. Half have implemented controls specifically to prevent it.
The consequences are measurable. Some 58 percent of enterprises have declined, delayed, or scaled back an AI initiative because of concerns about sensitive data leaving their premises or jurisdiction. For regulated industries, the constraint is not just risk appetite; it is regulatory reality. Financial institutions cannot send customer transaction records or proprietary trading algorithms to a public cloud AI service without risking violations of GDPR, PCI-DSS, or regional banking regulations. Insurers face equivalent exposure when processing policyholder medical records or actuarial models externally.
"Enterprises want to run AI where their data already lives," says Peter Sjoberg, VP of Solution Architects at Cloudian. "The challenge is that building a production-ready AI environment, with GPUs, storage, networking, and software frameworks, typically requires months of complex integration work. That is itself a barrier to sovereign AI adoption."
Glogowski sees the sovereignty imperative as a market signal NVIDIA has responded to directly. "Sovereignty isn't a niche requirement anymore – it's a tier-one enterprise concern. AIDP's answer is to run AI on the data, in place, on infrastructure the customer owns. The GPU is already next to the data. Sensitive content never has to leave to be processed, embedded, governed, or retrieved."
The economics of cloud AI are harder to predict than they appear
Cost has become a second and growing pressure point. The survey found that 40 percent of organizations report actual cloud AI spending exceeding initial projections, with 35 percent running between 10 and 30 percent over budget. Consumption-based pricing that fluctuates with data volumes and usage patterns was cited by nearly half of respondents as a barrier to expanding AI adoption.
For Sjoberg, the economic argument and the sovereignty argument reinforce each other. "When customers start modeling what it actually costs to run production AI workloads in the cloud at scale, the numbers often surprise them. And when you combine cost unpredictability with data governance exposure, the on-premises case becomes very straightforward, very quickly."
From components to capability: the integrated platform argument
The conventional approach to on-premises AI carries a risk that rarely surfaces in initial project plans. Assembling discrete infrastructure components into a validated production pipeline demands engineering expertise that most enterprises do not maintain in-house. Performance gaps tend to appear only under real workloads. Stack components that passed early testing prove incompatible at scale. And a custom-built environment, once deployed, becomes the organization’s ongoing responsibility to maintain, update, and re-validate with each new model or framework release.
The answer NVIDIA developed is NVIDIA AI Enterprise, a cloud-native software suite that provides the AI frameworks, inference microservices, and development tools enterprises need to build and run production AI workloads. It is designed to be deployable anywhere, including fully on-premises environments, and forms the software foundation on which validated enterprise AI platforms can be built.
Cloudian's HyperScale AI Data Platform (AIDP) builds directly on NVIDIA AI Enterprise, combining it with NVIDIA GPU infrastructure and Cloudian HyperStore S3-native object storage in a turnkey system engineered for enterprise environments. The platform deploys in days rather than months and requires no specialized AI skills to operate.
"Most AI delays aren't caused by models. They're caused by infrastructure complexity," says Glogowski. "AIDP removes a lot of the guesswork around what actually works at scale. Customers don't want to rebuild AI infrastructure from scratch every time. They want a validated architecture they can rely on in production."
The architecture matters as much as the components. HyperScale AIDP employs RDMA for S3-compatible storage technology, enabling direct data movement from storage to GPU memory with high sustained throughput and minimal latency. For workloads like real-time fraud detection, where transaction scoring must complete in milliseconds, or retrieval-augmented generation (RAG) applications where inference quality depends on continuously refreshing a model's context with proprietary organizational data, that pipeline performance is mission-critical.
"Storage is the part of the AI stack that doesn't get enough attention," says Glogowski. "But when you're running inference at scale, the speed at which data moves from storage to GPU is often what determines whether a use case is viable in production. RDMA for S3 was specifically designed to close that gap for object storage environments."
The platform ships with NVIDIA AI Blueprints – pre-built workflows for proven enterprise use cases. Financial institutions can deploy document intelligence against regulatory filings and internal policy libraries without exposing data externally. Insurance organizations can run video search and summarization across claims documentation and fraud investigation footage. Use cases that once required cloud exposure can now be executed entirely within the enterprise perimeter.
The hybrid horizon
The direction of enterprise AI infrastructure is not toward a wholesale rejection of cloud, but toward a more deliberate allocation of workloads. Over half of survey respondents plan to adopt or expand a hybrid approach with increased on-premises capacity over the next 24 months. Cloud retains advantages for development environments, burst compute, and AI use cases that do not touch sensitive data. On-premises infrastructure delivers the sovereignty, cost predictability, and low-latency performance that production AI increasingly demands.
"The question organizations should be asking is not whether to use cloud or on-premises, but which workloads belong where," says Sjoberg. "For AI that touches your most sensitive data, your most time-critical applications, and your highest-value proprietary information, the answer is increasingly clear: it belongs on your own infrastructure."
Glogowski sees the same bifurcation taking hold across NVIDIA's enterprise customer base. "The hybrid model is maturing. Organizations have gotten sophisticated about workload placement. What they want now is infrastructure that performs reliably at both ends – and a software stack that's consistent regardless of where it runs."
With 86 percent of enterprises expecting AI budgets to grow in 2026, the investment conditions are in place. What separates organizations that convert that budget into durable AI capability from those that struggle with complexity, cost overruns, and stalled initiatives is not the quality of the models they choose. It is whether the data feeding those models is rich enough, fast enough, and sovereign enough to make them genuinely intelligent.
Explore the survey “Data sovereignty, cost, and performance: The growing case for on-premises AI infrastructure” here.
Comments