NetApp introduced the Novus storage architecture with aggregate throughput above 100TB/s to remove metadata bottlenecks that limit large GPU clusters and AI agents. Its Data Director scales metadata separately from stored data, while the AI Data Engine discovers, classifies and vectorizes enterprise files. The move matters because many companies have models and applications ready but still fail to earn returns from AI investments.

NetApp Novus targets metadata bottleneck in AI data infrastructure

How Novus rebuilds storage for AI workloads

Novus was presented at NetApp INSIGHT in interviews with Chief Product Officer Syam Nair and Vice President of Product Management for Novus Gunna Marripudi to theCUBE Research hosts Christophe Bertrand and Rebecca Knight. The central claim is unified storage spanning on-premises and other infrastructure, instead of fragmented systems with manually applied governance policies. NetApp positions this unified layer as the condition for production AI, where data must be accessible, protected and ready for action.

Data Director is the core mechanical change inside Novus. It manages metadata apart from the data itself, so each layer can scale independently and concurrent metadata access supports many parallel requests from training and inference. Applications receive a unified view of files across multiple ONTAP storage clusters, without opening each cluster separately. According to Marripudi, concurrency of metadata access forms the pillar of the architecture.

Speed alone is treated as necessary but not sufficient for agents. The AI Data Engine adds discovery, classification and vectorization of enterprise data, with the stated aim of reducing preparation work that usually requires six to nine months of engineering. Nair argued that fast access to unprepared or unprotected data produces compromised answers rather than useful agent actions. The logic is that throughput must be paired with readiness and protection before models operate on business files.

What Novus changes for enterprise AI adoption

For companies deploying models and agents, the practical effect would be shorter data preparation and fewer separate storage silos to maintain. A team running several ONTAP clusters could expose files through one view, while GPU jobs gain parallel access without waiting on a single metadata path. Larger organizations with distributed infrastructure gain more from this consolidation, while smaller firms gain a simpler route to connect existing storage to AI workloads.

The second consequence concerns integration risk and vendor choice. NetApp points to open standards and the planned acquisition of PEAK: AIO Ltd. for parallel file system technology based on the pNFS protocol. Any Linux kernel released after 2018 already includes a pNFS client, which lowers the need to replace current systems. Buyers should still verify acquisition timing, supported configurations, performance on their file mix, and how classification and access controls map to their governance rules.

The marker to watch is completion of the PEAK: AIO acquisition and evidence that Novus sustains high concurrent throughput on live enterprise data, not only in benchmarks. Published deployment details and preparation-time reductions will show whether separate metadata scaling translates into faster production agents. If those cases appear, unified data infrastructure will become a stronger selection criterion for AI projects.