Akamai acquires LayerX, delivering end-to-end security and real-time AI usage control to any browser. Get details
Background

Architecting for Agents

Moving beyond hypercentralized infrastructure to meet agents where they actually live

Data from The State of AI Inference report reveals the architecture gap holding enterprises back

of organizations need end-to-end responses under 500 ms for their critical AI use cases

of organizations are still anchored to a single centralized cloud region

of total AI spend in production is on inference, not training

The edge of intelligence: Why centralized clouds can’t scale in the agentic AI era

Centralized clouds can’t scale for the internet of agents. Akamai exposes the wall facing legacy architecture, from stacking latency and hidden data-movement fees to shifting compute bottlenecks. Discover why real-time machine intelligence requires a fluid core-to-edge continuum.

Experts on the shift to distributed AI

Hear Akamai experts explain why AI workloads are moving out of centralized regions and closer to users, devices, and agents.

Learn the fundamentals. What’s driving the shift to distributed AI?

Experts explain how infrastructure will continue to act as the backbone of the AI era and what it takes to power true real-time intelligence at the edge.

Why do centralized data centers fall short for agentic AI?

Learn how scalable, distributed AI infrastructure forms the backbone of the AI era, powering workloads from the data center to the edge.

Why can’t brute-force, centralized build-outs scale to meet AI demand?

Akamai CTO Robert Blumofe explains why the shift from training to ubiquitous agents breaks the centralized-only model, and why an agent needs a hybrid of GPU, CPU, and storage placed across a core-to-edge continuum.

What does inference really cost once you model the whole bill?

Ari Weil, VP of Product Marketing and Cloud Strategy, breaks down the costs that sit outside the token price: the transfer fees charged every time a request crosses a zone or a cloud boundary, and the compute you pay for twice when you replicate a model to solve a distance problem.

Which metrics actually decide agentic AI performance in production?

Jon Alexander, SVP of Product for the Cloud Technology Group, explains why up to 90% of an agentic task runs off the GPU entirely, in tool calls and external API requests, and why the placement of that CPU work decides end-to-end latency.

How does a compute continuum work from core to edge?

Jon Alexander maps the continuum in practice: Large models stay in centralized clusters, agents and their tools run close to users, and an orchestration layer schedules the handoffs between them. He also lays out a crawl-walk-run path for teams already committed to a single region.

New AI survey: Inference breaks the latency wall

The State of AI Inference

See why the gap between real-time AI requirements and centralized infrastructure is pushing teams toward distributed inference.

Let’s talk AI

Fill out the form to:

  • See where latency and egress are quietly costing you today
  • Learn how to route your requests to the right tier
  • Explore which workloads move to the edge first