-
Centralized cloud architectures assume linear scaling, but high GPU utilization creates non-linear queuing delays that degrade Time to First Token (TTFT).
-
Practitioner survey data shows a major maturity gap: only 30% of early-stage AI experimenters prioritize proximity, jumping to 77% once workloads enter core production.
-
Heavyweight centralized "AI factories" excel at model training, whereas distributed air-cooled edge infrastructure is required for low-latency inference.
Video transcript
Video transcript
The interesting thing is that with all this AI, whether we are looking at inferencing, or I mean, that's what we focus on in our training models. Cloud is going to play, not going to play, it is actually the foundation of it. Can you talk about where does the traditional centralized cloud model start to break down for inference-heavy workloads, where we also need to look at decentralized approach? Also, we hear a lot about new clouds especially in Europe, because we can also throw the whole AI sovereignty, digital sovereignty that is going on in Europe a lot, but it will also become a global phenomenon.
Going back to the survey results that we saw, 60% of the practitioners that we surveyed said that proximity to the user is critical. But we also saw that in 46% of cases, their inference workloads are still running in a single centralized cloud region. And I think that gap is where we are starting to see people really learning some tough lessons about how to architect across Neo clouds, hyperscalers, alternative clouds, you know, and the rest of the folks who are competing in the ecosystem. And I think this is also an area where a lot of the tier one analysts are trying to really catch up and adjust their taxonomies.
Because if we look at data by maturity stage, then people in their early experimentation, only about 30% really identify proximity as critical. But when it's your core business workload, 77% of companies said that that proximity is critical. And I think that really speaks to when you start figuring out the business impact of the workload that you're building, and where you actually rely on that sort of round trip time, that's where people are starting to figure out unfortunately too late in the process.
I mentioned before, you know, when you do your first proof of concept, or proof of value, or whatever you consider it, that's when you need to start building in the infrastructure and the architectural relevance of where you are serving that POC. Because if we look at the people who are running single region deployments, and I said 77% of those workloads, they realized that proximity was important. Under 14% of those workloads are deployed in a centralized region or a single region. And I think that just shows you when you get down to over 85% of workloads being distributed, when they realize that latency matters and proximity is an important part of that latency equation, then you have a much clearer picture on trajectory.
And then you start asking yourself, where are those businesses located? If you thought that they were located in areas where, for example, sovereignty and privacy are top bill items like across EMEA, that would be a logical conclusion to draw. But it's not the case. The case is that the majority of those businesses are located in North America, specifically the United States, or they're in places in Asia Pacific, primarily in places like China and India and Japan, where we see a lot more scaling and more aggressive scaling of these inference workload.
And so I think for us, we're realizing that the centralized cloud was built for an era when compute was something that you went to. You had to go where the people had cards, horsepower, storage, et cetera. In the agentic era, we need something different. We need companies that are going to bring the compute closer to you. And more and more, not just users, but also people who are architecting their next applications are going to start being a lot more selective about picking the right card and the right infrastructure for their workloads and putting that in the right place for the users that they have to deploy to.
That's a completely different set of calculus than you used to apply in a centralized cloud region, because now it's not so much about just reserved instance capacity and committed revenue. Now it's really thinking about where and when do I have to scale, which is going to move us, I think, to a lot more of a just-in-time sort of deployment architecture than we've had in the last 20 years.
Do you feel that, you know, as, of course, AI goes more and more into production, it is already in production, do you see that AI infrastructure will become more hybrid with training will be centralized, but inference will be more distributed, as you also mentioned, that they do want it to be closer to user, but it's still centralized far from them. So how do you see training versus inference where they'll run?
I think the interesting thing is where we look at tokenomics and some of the companies that are really starting to push that idea forward and to give us a lot more insight into how a workload needs to align to a certain spec of machine, to a certain amount of, for example, RAM that's available to GPU and CPU processing. The thing that's happening is we need to better understand that scaling isn't always linear.
So centralized inference architectures always are going to assume that latency will roughly stay the same as the load increases, and that doesn't actually exist, that doesn't occur. As the GPU utilization is going to climb and you start thinking about saturating the available GPUs in a given location, then your queuing delays start to get exponentially larger and batching decisions that used to work at low loads start adding tens or hundreds of milliseconds to that time to first token, because your response time is just going to start shooting off the more that you have in queue and the more that you're going to be sensitive to either adding additional GPUs to the footprint that you have or having GPUs that might not be close enough to one another because you didn't plan for the appropriate size of GPU cluster and you're now routing away from where some of your requests are coming.
So you start thinking about, I've got a scaling dimension for how much I can scale up my GPUs, then I have a challenge around geographic dimensions. So you think about round trip times now being exacerbated by queuing times where I have to get access to resources in my centralized deployments and you start realizing that your load is going to reveal the sort of an architecture that you chose.
And so I think if we put all of this stuff together, the really interesting challenge is going to be how different cloud providers, whether you're a NeoCloud that's primarily focusing on a lot of the Tensor Core architectures, the GPUs that you need to drive sort of AI workload, but classically it would have been more of the training and less of the inference workloads. And then we see everybody diversifying and starting to add more of the cards that you need for inference. And then we start seeing these sort of hybrid or alternative types of clouds pop up. And you can't use either word because they already mean something in the cloud space.
But you have people that don't have as many locations, they might not have as much hardware, but they're very specialized for a given domain. Now you've got this really challenging architectural problem of how do I start planning my capacity based on available providers and help my providers understand how they should be building out their next infrastructure buys?
Because if I'm a company like Akamai, I've spent the last 28 years building out a very broadly geographically distributed network of relatively small data centers that I can now use primarily for inference workloads because they lend themselves to it. They're an air-cooled architecture, I don't need the same size cabinets, I don't need as much power, I don't need as much cooling. And so I will be able to build out a lot of globally distributed inference architecture.
If I'm a hyperscaler and somebody like an OCI for example or Oracle, they are now investing and really doubling and tripling down on a centralized infrastructure where they want to build out AI factories with very heavyweight GPUs like H100s, H200s, L40s and things of that nature that are specifically tuned for batch type of workloads.
Well how do I start going from my centralized batch to my distributed inference when I'm a business that's trying to scale my AI? And the answer is you start to really diversify your vendor pipeline and think about who's going to be providing you with access to compute and GPU and maybe in the future TPU architectures that you're going to need for your workload, where those workloads are being deployed and then you start to really think about your roadmap and your partners based on where they're going versus where they are because you need the ability to steer them in the direction that is going to be consistent with the workload and the architecture that you need.
And at this point there are some very large and very strategic companies that are steering this path for a lot of us out in the marketplace because they have the capital, the forethought and the maturity to already be steering people in a given direction but I think that is going to be the new arms race. It's not going to be centralized data centers, it's not going to be raw power or cooling resources that we've been seeing for you know call it the last three years very intensively.
Now it's going to be what is the specialization you need, who's got the network and the ability to deploy that specialization based on where you need it and how can you partner with them early enough and be meaningful enough in their roadmaps to ensure that their delivery roadmap aligns to yours.
That is a really interesting challenge that we're currently facing in 2026 going into 2027 and beyond but I think by the time we get to 2030 it will have completely have reshaped what the cloud ecosystem looks like with some new specialized providers really emerging out of the current set because they will have evolved to treat these specialized needs of distributed low latency inference workloads.
Frequently Asked Questions (FAQ)
Frequently Asked Questions (FAQ)
Centralized clouds force inference traffic through single regions, creating non-linear GPU queuing delays and batching bottlenecks when utilization spikes.
Time to First Token (TTFT) is a latency metric measuring the duration from when a user submits a prompt to when the model generates its initial response token.
AI training requires centralized, high-density, liquid-cooled compute clusters ("AI factories"), while AI inference requires geographically distributed, low-latency edge nodes.
Proximity minimizes round-trip network latency, preventing queuing bottlenecks and ensuring real-time response times for production AI applications.
Akamai Cloud leverages a globally distributed edge network equipped with air-cooled compute and GPU instances positioned close to end users and data sources.