Duality Technologies, a leader in privacy-enhancing technologies and secure data collaboration, has announced expanded support for Google Cloud’s Confidential Computing portfolio, now including confidential virtual machines powered by NVIDIA GPUs. This advancement enables large-scale secure AI workloads such as LLM training, encrypted inference, and privacy-preserving Retrieval-Augmented Generation (RAG).

With this upgrade, the Duality Platform now supports GPU-accelerated LLM inference and encrypted RAG within trusted execution environments (TEEs)—a major performance improvement over previous CPU-only capabilities.

Customers can run end-to-end secure generative AI workflows on NVIDIA GPUs using Google Cloud Confidential Computing. By combining full-stack data confidentiality with the high performance of NVIDIA H100 GPUs, organizations can execute confidential AI workloads that were previously too slow or impractical due to latency and throughput limitations.

“This changes the game,” said Dr. Alon Kaufman, CEO and Co-Founder of Duality Technologies. “Our customers can now run privacy-preserving AI with LLMs at production scale. With GPU acceleration, the performance bottlenecks of secure computing are gone—making secure LLM training and inference practical.”

The capability is powered by Google Cloud’s Confidential Space and confidential VMs using NVIDIA H100 GPUs, with integrated support for Intel TDX and Cloud KMS. Duality successfully validated this technology by running a Mistral-7B model with encrypted vector RAG (via Faiss) inside a fully confidential pipeline.

“With Confidential GPUs, organizations can process sensitive AI workloads entirely within trusted execution environments without giving up performance,” said Nelly Porter, Director of Product Management at Google Cloud. “Pairing NVIDIA H100-powered confidential VMs with Duality’s encrypted workflows enables secure LLM training and inference at scale, with end-to-end protection from data leakage.”

Key Highlights

  • GPU Support for Confidential AI: Secure LLM training, inference, and encrypted RAG on confidential NVIDIA H100 GPUs.
  • Scalable Performance: Dramatically faster runtimes compared to CPU-only secure workloads.
  • Enterprise-Ready: Supports regulated industries including defense, healthcare, and finance.
  • Seamless Cloud Integration: Available through Dynamic Workload Scheduler in Google Cloud's Confidential Space.

Previously, confidential computing was constrained to CPU-based environments—adequate for testing but insufficient for enterprise-scale AI. With confidential GPUs now part of Google Cloud’s security stack, Duality customers can run both training and inference securely inside TEEs, enabling high-throughput, privacy-preserving AI workflows across sectors.

This capability is initially available on Google Cloud’s Confidential A3 VM type in preview, with expanded availability expected later this year.

For more insights on AI, IoT, cybersecurity, and emerging tech trends, visit Itech360hub.com.