The AI ??industry's focus is shifting from large-scale model training to routine inference, making CPUs with massive addressable memory capacity a market necessity. At its annual Discover conference, Hewlett Packard Enterprise (HPE) officially announced that it will integrate the ProLiant DL394 compute server, powered by NVIDIA Vera CPUs, into its Private Cloud AI solution. At the same conference, HPE also released a series of updates to its data lake, network equipment, and other product lines.
Two years ago, HPE launched Private Cloud AI, delivering an integrated AI hardware and software stack for enterprises. The complete solution integrates HPE servers, storage, networking, and supporting software, and is built entirely on the official reference architecture of NVIDIA GPUs, Ethernet platforms, data processing units (DPUs), and network interface cards.
HPE Executive Vice President and Chief Technology Officer Fidelma Russo stated that the new version of the Private Cloud AI inference cluster can be scaled up to a maximum of 256 Blackwell architecture GPUs; it is also officially compatible with NVIDIA's new generation Arm architecture Vera CPU released at GTC in March this year.
HPE DL394 Gen12 Inference Server
HPE debuted the DL394 Gen12 server at Computex Taipei this month, becoming the core hardware for the private cloud AI system supporting Vera CPUs.
The entire machine is based on an NVIDIA Vera CPU, featuring 88 custom Olympus Arm cores and 176 threads; a single unit supports up to 3TB of LPDDR5X memory, with a CPU memory bandwidth of up to 1.2TB/s. The entire machine adopts a 2U air-cooled chassis design, specifically optimized for AI inference workloads.
In addition to the newly added DL394 inference server, HPE's private cloud AI suite has simultaneously launched several security and management capabilities: It integrates HPE Zerto security software to defend against malicious AI agent attacks, while providing continuous data protection and data time-lapse capabilities; it also supports permission isolation and behavior restrictions for locally running AI agents.
HPE has also upgraded two supporting AI delivery solutions: AI Factory at Scale and Sovereign AI Factory, adding support for NVIDIA Confidential Computing to ensure the security of models and privacy data throughout their entire lifecycle in local deployments and sovereign compliance scenarios.
This solution builds a trusted security chain through encrypted authentication links and combines BlueField DPU and the DOCA network development kit (known in the industry as "CUDA for the network domain") to achieve end-to-end encrypted isolation.
Data Architecture Comprehensive Upgrade, Integrating MCP with Open-Source Scheduling Tools
HPE Private Cloud AI's Data Fabric has undergone a major iteration: While maintaining the original native compatibility architecture of the HPE Data Fabric software, this update extends the Model Context Protocol (MCP) to adapt to the open-source data flow scheduling tool Apache Airflow. This standardized protocol allows for batch mounting of metadata to distributed data, enriching the context information readable by AI agents.
Russo added that deploying Data Fabric software is challenging for enterprise customers. Therefore, HPE has launched a dedicated ProLiant appliance deployment option, with Data Fabric components pre-integrated, significantly shortening the deployment cycle and simplifying maintenance.
High-Performance Storage Support, 20x Improved Inference Response Speed
Alletra Storage MP X10000 distributed storage has long been a standard feature in private cloud AI. When this storage is connected to a DL380a Gen12 server equipped with eight NVIDIA H200 GPUs, large model token response latency is reduced by up to 20 times.
Russo explained that the X10000 became NVIDIA's first globally certified storage array adapted for object storage AI workloads this spring, and now it adds dual-protocol support for file storage.
"Our storage architecture differs from industry competitors. There is no performance trade-off between object storage and file storage; both types of storage run in parallel, perfectly adapting to mixed AI analytics workloads."
The entire machine incorporates a full suite of metadata management, polling acceleration, metadata enhancement, and key-value caching acceleration services, significantly improving overall computing throughput and inference performance compared to previous generation private cloud AI solutions.