Akamai acquires LayerX, delivering end-to-end security and real-time AI usage control to any browser. Get details
Background

GPU on Akamai Cloud

Turn vision into reality — power your ambitious AI strategies with NVIDIA GPUs

Deploy the right cloud GPU for your workload and budget

Sustainably test and scale compute-intensive workloads through Akamai’s GPU cloud services without draining your infrastructure budget. Akamai Cloud lets you choose the right GPU for your workload. Easily deploy cloud-based GPUs when needed. Predictable pricing and low-cost egress help you manage costs and scale machine learning, AI workloads, data processing, and other high-performance computing with confidence.

Powering your AI strategy with Akamai GPUs

Power AI inference with globally distributed GPU compute designed to deliver low-latency responses at scale.

Scale massive LLMs

Scale private inference across large models with NVIDIA RTX PRO™ 6000 Blackwell GPUs for compute-intensive AI.

Perform high-density streaming and transcoding

The NVIDIA RTX™ 4000 Ada GPU powers live 8K transcoding and AI upscaling with high throughput at the edge.

Create professional 3D models and visualizations

Accelerate CAD workflows, server-side rendering, and interactive 3D modeling for enterprise apps with NVIDIA RTX™ 6000 Quadro GPUs.

Using Akamai GPUs to power your vision

Select

Choose from NVIDIA RTX PRO™ 6000 Blackwell or Ada architectures, matching your performance requirements to meet your budget.

Configure

Customize compute, memory, and storage to optimize for your unique workload.

Deploy

Launch GPU instances in minutes across Akamai’s cloud locations, meeting your user and data wherever they are.

Realize

Achieve your ambitious AI goals with high-performance NVIDIA compute designed for immersion, autonomy, and scale.

Why Akamai GPUs

Dedicated GPU resources

Get strong, predictable performance with dedicated GPU resources.

25x media encoding performance

RTX™ 4000 Ada GPU plans are built for media workloads, with 2x encoding, 2x decoding, and 1x AV1 encode/decode engines per card.

Superior AI inference performance

Deliver up to 60% lower latency and 3x higher throughput, for up to 86% less cost with image generation and AI workloads compared to equivalent hyperscaler GPUs.

Features

  • Low, predictable GPU pricing
  • Add GPU nodes to managed Kubernetes clusters with LKE
  • Manage infrastructure flexibly with our UI, API, CLI, and developer tool integrations
  • Save up to 90% on egress with Akamai Cloud (US$0.005 per GB in most regions)
  • Set up CI/CD pipelines with custom images, Terraform provider, and more
  • Access full documentation support to install NVIDIA CUDA toolkit to get started
  • Easily resize a Shared CPU or Dedicated CPU instance to a GPU instance
  • Configure Backups to retain data with automated daily, weekly, and biweekly snapshots
  • 24/7/365 email and phone support for all customers

Akamai GPU Use Cases

Agentic and multimodal AI workloads

The shift from reactive chatbots to autonomous agents capable of multistep reasoning and independent task execution requires agentic AI. Agentic AI requires specialized hardware to support the massive throughput needed to simultaneously process text, visuals, and audio to maintain the real time responses that users expect.

NVIDIA RTX PRO™ 6000 Blackwell provides 96 GB of VRAM per GPU in a high-throughput architecture to mitigate “bottlenecking” found in shared cloud resources or legacy architectures.

By moving these memory-intensive workloads to Akamai’s edge, customers are able to minimize latency in backhauling data to centralized data centers and provide an efficient multimodal experience to their users.

Real-time conversational AI

Conversations move at the speed of human thought. For AI to join that conversation in a truly immersive way, the AI needs to be able to be able to keep up. Real-time conversational AI is defined by its ability to process, reason, and respond within that line of thought. When this happens, the digital interface goes from feeling like an input to a natural and fluid interaction with AI.

Akamai Cloud has the GPUs to take this conversation global, scaling out while maintaining the predictable economics and industry-leading egress rates. Whether you are building a retail concierge to guide a purchase or creating an interactive educational tutor, our GPUs empower you against distance and unpredictable costs. With Akamai, you have the ability to turn interactions into rapid engagement.

Physical AI and computer vision

Physical AI marks the transition from digital intelligence to real-world action. To turn this vision into a reality, AI must live where the data is born: at the edge. Whether it’s managing crowd safety in public transport, optimizing store inventory in real time, or automating quality control on a factory floor, physical AI and computer vision require instantaneous processing of massive video and sensor streams. Sending this data back to a centralized data center isn’t just expensive; the latency makes real-time decision-making not real time.

The NVIDIA RTX™ 4000 Ada is an optimal engine for this transition. Designed for high-density efficiency, it allows you to process multiple high-resolution camera feeds simultaneously with the throughput needed for advanced object detection, segmentation, and spatial reasoning.

By deploying these GPUs on Akamai’s cloud, you bring intelligence to physical sites. This proximity allows autonomous systems or safety monitoring tools to rapidly react. Akamai’s distribution also enables you to address critical challenges like data privacy and sovereignty. By configuring personally Identifiable Information (PII) and sensitive telemetry for local processing, you can support your security posture while reducing costs associated with backhauling massive datasets to a central data center. Akamai helps achieve optimal price-performance not only with highly performant GPUs at a reasonable price, but also by putting you in control of latency and egress costs. Take your AI out of the data center, and put it to work in the real world.

Live video transcoding

Standard definition is anything but standard. With today’s 8K videos and high-fidelity media, you need a transcoding engine that doesn’t just convert formats but enhances every frame of that beautiful 60 FPS live video. Transcoding at the edge allows you to deliver the “live in HD” experience that your customers want for all of their sports, cat, or viral video needs without the crippling latency or massive costs of data streams back to a centralized data center.

The NVIDIA RTX PRO™ 6000 Blackwell GPU has the VRAM to support transcoding for a wide variety of 8K video formats. This hardware allows you to configure for live transcoding alongside AI upscaling, automated object detection, and dynamic ad insertion within a single unified workflow. Leveraging dedicated hardware encoders enables you to process videos several times faster than CPU-based methods of the past, allowing your live streams to minimize buffering for millions of concurrent viewers.

Akamai’s cost-effective cloud prices mean that our Blackwell GPUs can achieve incredible price-performance while reducing the overall cost of video encoding.

Rendering and simulation

Ambitious creative and scientific visions require more than just raw compute; these workloads require a platform that can break through the “memory wall.” High-fidelity rendering, complex industrial simulations, and large-scale digital twins are among the most resource-intensive workloads in existence. To execute these without compromise, infrastructure must provide massive parallel processing power and extensive VRAM to handle the heavy geometry and high-resolution textures that modern professional applications demand.

The NVIDIA RTX™ 4000 Ada, deployed on Akamai’s globally distributed cloud, is a definitive engine for this strategic scale. Featuring power-efficient Ada Lovelace architecture and 20 GB or high-speed GDDR6 memory, this architecture allows engineers, architects, and VFX artists to load massive datasets into GPU memory, allowing you to have control over data-swapping schema.

Whether you are running real-time ray tracing for cinematic visualization or executing complex fluid dynamics and molecular modeling, Akamai Cloud GPUs provide the stability and professional-grade performance needed to support the move from concept to final output at today’s speed.

Akamai’s predictable, transparent pricing and low egress fees allow you to move massive simulation results or high-resolution frames without budget surprises. By leveraging Akamai’s edge native distribution, you can also support globally distributed teams, delivering high-fidelity visualization and simulation outputs locally for a seamless, collaborative experience.

High-fidelity visualization

 

With the globalization of talented employees, there is an increasing number of globally remote design and engineering teams. These distributed teams require the ability to visualize complex data in high fidelity in order to do their job. High-fidelity visualizations, like immersive 3D architectural walk-throughs or real-time product configurators, depend on pixel-level details and near–zero latency interactions. Traditional centralized hosting models fail to provide the required experience for these cross-continent teams because the distance introduces too much lag, destroying the immersive experience and stalling productivity.

The NVIDIA Quadro® architecture, hosted on Akamai Cloud, is designed to turn these professional-grade visions into reality. Unlike standard hardware, Quadro GPUs provide the specialized driver stability and memory capacity needed to support ISV-certified applications in CAD, BIM, and engineering.

By placing these resources physically closer to your designers and stakeholders, Akamai allows you to deliver workstation-class performance with worldwide reach, without the need for expensive local hardware.

Empower your global teams to collaborate in real time on your ambitious projects. With Akamai, you can reduce the impact of distance on your visualization strategy, allowing you to move from complex datasets to stunning, interactive realities with unprecedented agility and scale.

Explore more about Akamai compute

Accelerated Compute

Optimize performance with ASICs and NETINT VPUs for faster media transcoding.

CPU

Match Shared CPU, Dedicated CPU, or High Memory compute plans to your applications’ needs and budget.

Frequently Asked Questions (FAQ)

Frequently Asked Questions (FAQ)

A cloud GPU, or graphics processing unit, is a specialized hardware component designed to accelerate tasks that require significant parallel processing power. See details.

Akamai Cloud offers NVIDIA RTX PRO™ 6000 Blackwell, NVIDIA RTX™ 4000 ADA, and NVIDIA Quadro® GPUs. See NVIDIA GPUs.

Akamai recommends GPUs based on your workload type and scale:

  • RTX PRO™ 6000 Blackwell – large AI inference, NLP, vision, video analytics, high-end rendering
  • RTX™ 4000 Ada – smaller AI models, graphics, rendering, video workflows
  • NVIDIA® RTX™ 6000 Quadro – CAD and professional visualization

Choosing the right GPU maximizes performance and cost efficiency without overprovisioning. Explore more or request access.

GPUs are available in select Akamai Cloud regions. Akamai operates a globally distributed cloud with 25+ core regions and 4,400+ points of presence worldwide, and our team can help identify the best GPU location based on your workload and latency needs.

Yes. Akamai offers NVIDIA RTX PRO™ 6000 Blackwell Server Edition® GPUs in 1-card, 2-card, 4-card, and 8-card plans so teams can choose the right level of memory and throughput for their workload — from fine-tuning and recommendation engines to large multimodal inference, 8K video, and gaming.

Request GPU access to connect to the team. For quickstart, check out our GitHub repo to deploy an open source LLM on Kubernetes.

Akamai is constantly benchmarking compute performance to help customers get the price-to-performance ratio needed. The latest published benchmarks show the NVIDIA RTX PRO™ 6000 Blackwell running on Akamai Cloud delivers up to 1.63x higher inference throughput than the H100, achieving 24,240 TPS per server at 100 concurrent requests.

Akamai offers flexible GPU purchasing options that vary by GPU type, region, and deployment size. Based on your workload and scale needs, we provide the right mix of commitments, contract terms, GPU models, and shape sizes to fit your requirements.

Akamai’s GPU pricing is often more predictable than traditional cloud models, with clearer cost structures for compute and fewer hidden fees. Lower egress costs help reduce the expense of moving data, while workload-based planning allows teams to align usage with actual demand, improving cost efficiency for AI inference and other high-performance workloads.

A distributed cloud GPU reduces latency by running AI inference closer to users and data, making it ideal for real-time applications that require fast, consistent responses. It also improves efficiency by avoiding unnecessary data movement and enabling scalable performance across locations. Learn more about AI inference hardware decisions and distributed AI inference.

Akamai Cloud GPU instances are well suited for AI inference, live video transcoding, rendering, simulation, and visualization workloads that benefit from high-performance, parallel compute. These use cases often require both speed and scalability, especially when delivering real-time or visually intensive experiences. Explore how this supports media solutions for live video delivery.

Next steps