Akamai Sharpens Its AI Edge with Launch of Akamai Cloud Inference

Akamai (NASDAQ: AKAM), a leader in cybersecurity and cloud computing, has introduced Akamai Cloud Inference, a breakthrough service designed to enhance AI performance and efficiency. This new solution enables businesses to achieve a 3x improvement in throughput, 60% lower latency, and 86% cost reduction compared to traditional hyperscale infrastructure. Running on Akamai Cloud, the world’s most distributed platform, Cloud Inference addresses the growing challenges of centralized cloud models.

“Getting AI data closer to users and devices is hard, and it’s where legacy clouds struggle,” said Adam Karon, Chief Operating Officer and General Manager of Akamai’s Cloud Technology Group. “While training large language models (LLMs) will continue to occur in big hyperscale data centers, inference—the real-world application of AI—needs to happen at the edge. Akamai’s platform, built over two and a half decades, uniquely positions us to power this AI future.”

AI Inference on Akamai Cloud

Akamai Cloud Inference offers platform engineers and developers the tools to run AI applications and data-intensive workloads closer to users, improving performance and reducing costs. Key features include:

  • Compute: Akamai Cloud provides a diverse compute environment, from CPUs for fine-tuned inference to high-performance GPUs and tailored ASIC VPUs for a range of AI workloads. Integration with NVIDIA AI Enterprise optimizes inference using Triton, TAO Toolkit, TensorRT, and NVFlare.
  • Data Management: Akamai’s partnership with VAST Data enables real-time access to critical AI inference data. The platform supports scalable object storage and integrates with leading vector database vendors like Aiven and Milvus to enhance retrieval-augmented generation (RAG).
  • Containerization: Built on Kubernetes, Akamai Cloud Inference supports demand-based autoscaling, hybrid/multicloud portability, and enhanced resilience. Powered by the Linode Kubernetes Engine (LKE)-Enterprise, it seamlessly integrates open-source projects like KServe, Kubeflow, and SpinKube for efficient AI model deployment.
  • Edge Compute: Akamai AI Inference incorporates WebAssembly (Wasm) for lightweight, low-latency AI execution. Partnering with Fermyon, developers can deploy AI inference directly from serverless applications, optimizing real-time processing.

These capabilities enable businesses to build AI-powered applications with low latency and high scalability, ensuring a seamless user experience. Akamai Cloud Inference operates on Akamai’s globally distributed platform, spanning 4,200+ points of presence across 1,200+ networks in over 130 countries.

The Shift from Training to Inference

As enterprises mature in AI adoption, the focus is shifting from training LLMs to real-world inference. While LLMs excel at general tasks like summarization and customer service, they are costly and complex to train. In contrast, lightweight AI models optimized for industry-specific challenges provide greater efficiency, measurable outcomes, and a stronger return on investment.

The Need for a More Distributed Cloud

With increasing data generation outside centralized data centers, businesses require AI solutions that process information closer to its origin. Distributed cloud and edge architectures are emerging as the preferred choice for operational intelligence, delivering real-time insights across globally distributed assets. Early use cases on Akamai Cloud include:

  • In-car voice assistance
  • AI-powered crop management
  • Image optimization for e-commerce marketplaces
  • Virtual garment visualization for online shopping
  • Automated product description generation
  • Customer sentiment analysis

“Training an LLM is like creating a map—gathering data, analyzing terrain, and plotting routes. It’s slow and resource-intensive, but highly useful once built,” explained Karon. “AI inference, on the other hand, is like using GPS—applying knowledge instantly, recalculating in real time, and adapting to changes. Inference is the next frontier for AI.”