Background

중앙 집중식 AI가 대규모로 확장되지 못하는 이유

Akamai CTO 로버트 블루모프(Robert Blumofe) 박사가 어디서나 활용되는 AI 에이전트에 맞춰 중앙 집중식 인프라를 확장할 수 없는 이유를 설명합니다.

핵심 내용

추론은 엣지에서 실행해야 함

중앙 집중식 고밀도 GPU 데이터 센터는 AI 모델 학습에 적합하지만, 전 세계에서 추론을 실행하려면 실제 애플리케이션에 필요한 짧은 지연 시간과 높은 대역폭을 제공하는 분산형 엣지 아키텍처가 필요합니다.

에이전트는 여러 구성요소로 이루어진 시스템

진정한 AI 에이전트는 LLM이 추론과 의사 결정을 담당하고, 데이터베이스 쿼리, 검색, 이메일과 같은 특수 목적의 비 AI 툴이 실제 작업을 수행하는 복합 시스템입니다.

비 AI 툴과 더 작은 모델을 우선 활용

지속 가능하고 비용 효율적인 AI 엔지니어링을 위해 개발자는 각 작업에 적합한, 가장 작고 가장 특화된 모델을 사용해야 합니다. 또한 가능하면 더 저렴하고 신뢰할 수 있는 비 AI 툴을 사용하도록 에이전트를 설계해야 합니다.

"에이전틱"으로의 전환이 웹을 바꿀 것

웹이 인터넷을 바꾸어 놓았듯, 어디서나 이루어지는 AI 추론은 정적인 페이지 탐색과 링크를 멀티모달 기반의 대화 중심 에이전트 경험으로 대체하며 우리가 기술과 상호 작용하는 방식을 재정의할 것입니다.

타임스탬프가 포함된 요약

[00:00–01:46] 중앙 집중식 AI의 한계: 스왑닐 바르티야(Swapnil Bhartiya)는 실시간 AI 에이전트를 위해 중앙 집중식 클러스터에 의존하는 것의 문제점을 지적하며 이야기를 시작합니다. 로버트 블루모프(Robert Blumofe) 박사는 초기 모델 학습 단계에서는 중앙 집중식 GPU 인프라가 필수적이지만, 중앙 집중식 모델만으로는 "유비쿼터스 AI"(추론, AI 애플리케이션, 능동적 AI 에이전트로의 전환)를 지원할 만큼 확장할 수 없다고 설명합니다.

[01:46–03:45] 초기 웹과 사이버 보안의 역사적 유사성: 블루모프는 중앙 집중식 호스팅으로 인해 “월드 와이드 웨이트(World Wide Wait)”라는 현상이 발생했던 웹의 초기 시절과 사이버 보안의 발전 과정을 비교합니다. Akamai는 수학과 분산 시스템, 알고리즘을 활용해 웹의 확장성 문제를 해결하고, 엣지에서 애플리케이션을 제공하는 방법을 개발했으며, 이제 이 방법은 AI 추론 분야에도 적용되어야 할 필수적인 청사진이 되었습니다.

[03:45–05:20] 학습 및 추론 인프라 비교: 블루모프는 훈련과 추론의 서로 다른 아키텍처 요구사항을 물리적 “적합성”에 따라 분석합니다. 학습은 중앙 집중화된 데이터 세트와 고밀도 GPU 클러스터에 근접해야 하는 반면, 추론은 전 세계에 분산된 최종 사용자와 동적 멀티모달 상호작용에 근접해야 합니다.

[05:20–07:19] AI 에이전트의 하이브리드 아키텍처: 블루모프는 AI 에이전트가 여러 LLM과 비 AI 툴(예: 스토리지, CRM 접속, 벡터 데이터베이스, 웹 검색)로 구성된 복잡한 시스템이라고 설명합니다. 이러한 이유로 에이전트는 대규모 GPU 전용 클러스터 대신 GPU, CPU, 스토리지를 모두 갖춘 하이브리드 인프라를 필요로 합니다.

[07:19–08:34] Akamai AI 그리드 지능형 오케스트레이션: Akamai가 최근 4,000개 엣지 위치에 새로 출시한 AI 그리드에 대해 블루모프는 이 그리드가 제한적이고 중앙 집중화된 “일률적” 솔루션 대신 적절한 인프라(GPU, CPU, 스토리지의 맞춤형 조합)를 적절한 장소와 시간에 제공하는 것을 목표로 하고 있음을 설명합니다.

[08:34–10:04] 지속 가능한 AI 운영 달성: 블루모프는 급증하는 토큰 비용과 에너지 수요를 해결하기 위해 개발자들에게 특정 기법에 무작정 의존하는 방식에서 벗어나라고 조언합니다. 그는 더 작지만 고도로 특화된 모델을 활용하고 기본적인 작업에는 더 효율적인 비 AI 툴을 우선시함으로써 더 스마트한 에이전트를 개발할 것을 제안합니다.

[10:04–10:55] 클라우드 공급업체 종속성과 전환 경로: 블루모프는 이 업계가 아직 초기 단계에 있기 때문에 대규모 클라우드 공급업체에 대한 종속성이 아직은 큰 장애물이 아니라고 합니다. 대부분의 모델이 표준 OpenAI API를 사용하기 때문에 개발자는 공급업체를 쉽게 변경하고 점진적으로 분산형 모델을 도입할 수 있습니다.

[10:55–12:00] 향후 현실에 적용될 에이전틱 AI 응용 사례: 블루모프는 일상적인 디지털 경험이 어떻게 에이전트 기반 인터페이스로 변화할 것인지에 대해 설명합니다. 사용자는 링크를 일일이 복사하거나 정적인 웹 페이지를 클릭하는 대신, 맞춤형 비디오, 이미지, 툴을 실시간으로 제공하는 전문 AI 에이전트와 대화형으로 소통하게 됩니다.

[12:00–13:21] 엣지 컴퓨팅과 로컬 캐싱: 블루모프는 기존의 엣지 모델이 AI에 어떻게 적용될 수 있는지 살펴보며, 소규모 모델 실행, 보안 점검 수행, 최적의 백엔드 리소스를 판단하기 위한 신속한 라우팅 결정과 같은 작업이 엣지에서 바로 효율적으로 이루어질 수 있다고 전합니다.

비디오 자막

비디오 자막

**Swapnil Bhartiya:** AI conversations usually revolve around massive GPUs and centralized clusters. But real world AI agents demand millisecond response time at the point of contact, and centralization simply cannot scale to meet that moment. Akamai is rethinking AI cloud infrastructure by building a distributed grid for inference across 4,000 global locations. Joining us today once again is Dr. Robert Blumofe, EVP and CTO at Akamai. Robert, it's great to have you on the show.

**Dr. Robert Blumofe:** Thanks. Thanks for having me.

**Swapnil Bhartiya:** We are seeing massive investment in centralized AI data center right now. Why do you believe this centralized everything approach is the wrong model for the future of AI?

**Dr. Robert Blumofe:** So it's a great question and ultimately the central thesis is simply that the sort of brute force approach of building out large amounts of infrastructure in centralized locations ultimately, well, it's expensive. But ultimately even that expense aside, can't achieve the scale that's going to be needed as AI sort of moves into its next phase that you might characterize as ubiquitous AI. And I think it's worth maybe highlighting a couple of ways in which the demand has changed, because ultimately you need to look at the demand and see how the infrastructure aligns to that demand. And I would focus maybe on two shifts. One would be the shift from training to inference, and the other one I would characterize as the shift from sort of the early days of a chatbot to an AI application or an AI agent. You know, it wasn't that long ago focusing on the first of those shifts. It wasn't that long ago that most of the infrastructure demand really came from the training use case where you were. And by and large I'm talking about the pre training of large generative models like LLMs that was driving a huge amount of the infrastructure demand. And in that use case, absolutely centralized, large scale, dense GPU infrastructure makes a whole lot of sense. But as you move into inference, it changes a lot. And of course, and I think we all know this, that training is really sort of a mandatory cost that is necessary to realize the value through inference. All the value in AI comes from the inference. And training is simply an investment that we have to make to realize the return that you get through through inference. And now as we're moving into a more mature phase, much more of the demand is coming from inference. And that's a good thing because again, that's where we get the value. So inference driving demand is a very, is a very good thing. And I would argue that the nature of inference is changing quite a bit. And again, that's the shift I'm talking about from the chatbot to the AI application or the AI agent. You know, in the case of the chatbot, I think we really thought of AI as sort of a destination. It was intentional. You went to chatgpt.com to use AI or you fired up your anthropic Claude desktop to use AI. It was intentional. It was a destination. Once you move into AI powered applications and AI agents, it becomes ubiquitous. It's no longer a specific destination, it's just part of everything that you do. Certainly everything that you do online, you go to a website to look for a car AI, you go to a healthcare provider to make an appointment to see your doctor. AI, everything that you're doing is AI powered, probably even everything that you're doing on your desktop, even irrespective of the web, you know, you want to send something to, to your kids AI. So AI becomes ubiquitous, that changes the nature of the demand. And in that world where AI is ubiquitous, being used all the time by everyone, a centralized approach just really isn't, isn't going to cut it. And we risk sort of revisiting the old. You know, back then we called it the world wide wait. It could turn into large language molasses for lack of a better term.

**Swapnil Bhartiya:** If AI training is more or less like writing software, inference is like deploying it globally so people can use requires totally different infrastructure. You have handled high latency sensitive use cases like live sports and global cybersecurity. How do the lessons from those challenges apply to AI infrastructure today?

**Dr. Robert Blumofe:** Yeah, that's a great point. And I do see a lot of parallels to the early days of the web. Also, I think some of the changes that happened maybe a decade or so ago in cybersecurity. So in many ways people say this all the time. History does repeat itself and it's kind of repeating itself for the umpteenth time here. So while there's obvious differences, there's a lot about what's happening now that I think does parallel what we saw in the early days of the web. As the web was getting popular, really transforming the Internet, there were a lot of concerns that the web simply wouldn't scale to meet the demand. And in many ways those concerns probably were well founded because you did have a situation where web applications were centralized. Now we were pre cloud, but we did have hosting providers. And arguably the hosting providers back then were even more centralized than today's hyperscalers. By and large, most of the infrastructure was in the U.S. for example, was heavily located in places like Ashburn and San Jose. So every time you used a web application, you had to traverse large distances into a handful of centralized locations. And while that might have been okay in the very early days, where a website was a pretty static thing, just some text, maybe a few images, as you move into video, for example, and large demand for that video, you simply cannot meet the bandwidth requirements and the latency requirements through that centralized model. And that's really, I think, what led people to, you know, jokingly say that the World wide Web should be, you know, called the worldwide, worldwide wait. And people speculated that the web would simply collapse. And really that concern was the beginning of Akamai, where, you know, Tom and Danny, the two founders, came forward with a better approach, math algorithms, distributed systems, rather than brute force. And they showed that you can actually deliver websites and web applications from the edge of the Internet, dramatically increasing the available bandwidth, dramatically lowering the latency. And really that's what made the web work. And that was a critical ingredient. Also, as the web transitioned from these static sites to dynamic, where your communication is happening all the time, it's not just click on a link and wait for a response. You're constantly interacting with these web applications. CDNs made all of that work. And a similar approach, really the same approach, is in many ways what enabled powerful cybersecurity defenses. Because cybersecurity also went through a pretty strong transformation about 10 years ago, maybe a bit less. Where you moved from our biggest concern being things like Anonymous to sophisticated ransomware and the world of sophisticated attackers. With, with ransomware and powerful DDoS, extortion attacks, the centralized approach just wouldn't work. And again, you have to borrow from this playbook of math, distributed systems, algorithms, and that worked. AI today, I think, is in a very similar regime where as you move from training to inference, as you move from fairly low bandwidth and high latency types of interactions like the chatbot, where you're simply typing some text, waiting for a response type, typing some text, waiting for a response, you move from that into an AI powered application or an AI agent, the nature of the demand just changes dramatically. It becomes ubiquitous. It's constant, it's high bandwidth, it requires low latency. And the, again, the brute force approach just isn't going to work. You can't do this with purely centralized infrastructure. The demand has moved to the edge. So the infrastructure and capabilities of AI have to also move to the edge.

**Swapnil Bhartiya:** Let's look at training versus inference. Most companies care far more about inference to actually deliver AI to their users. Do these two faces genuinely demand completely different infrastructure architectures?

**Dr. Robert Blumofe:** Yeah, it's a great point and a great way to distinguish these use cases. And I think it's helpful to think about what is the important affinity. What does the infrastructure need in terms of proximity and, and arguably in the case of training, the key affinity, the key proximity requirements is the data set, the training data set. And most training data sets are fairly large and they're generally stored in some large storage cluster that's going to be fairly centralized. You typically wouldn't have your training data set distributed around a large number of locations. It's going to be fairly centralized, so it makes sense to do the training where the training data set is. Um, it's also the case that if you look at the actual computation, you know, it's very GPU dense. So a dense GPU cluster centralized where the training data set is, that makes a whole lot of sense when you move into inference. Well, what's the affinity? What does it need to be near? Well, it needs to be near the things that it's interacting with. And, and there's a lot of things that, that AI applications, AI agents have to interact with. But obviously one of the important users, of course, is the people. Us. You know, we are going to use agents to get things done for us. We're going to engage in conversations with these agents to help specify what it is we want done, to look at results, to review results, provide feedback. It's going to be very conversational. So the affinity of an agent, I mean, we can get into this a little bit more in a little bit because there's a lot of different affinities, but one of them clearly is to the users. And users are typically distributed over a fairly large swath of geography, whether it's a country or a continent or, or the entire world. So it makes no sense really for the agent that you and I are interacting with to be thousands of miles away, centralized in a single location. And relative also to what I was mentioning earlier about the nature of the interaction changing with the web, where we went to high definition video and things like that. The same thing is the case with agents. You don't want to think of an agent interaction as being simply text or even voice. A good agent is going to show us video, is going to show us images, and it's going to be dynamically updating the video and dynamically updating the images. These are capabilities that you don't get outside of AI. And it's one of the great benefits of using an AI agent is that you have all these modalities available, video, images, interaction that isn't available in with other technologies. So you know, as we see agents become more ubiquitous, I think we'll see these high bandwidth forms of interaction really take, take hold because they really deliver value and they deliver something compelling and interesting. And there's just no way to do that from through a brute force build out in centralized infrastructure.

**Swapnil Bhartiya:** When we talk about AI agents, everyone immediately fixates on GPU scarcity. Can you explain why AI agents actually require a hybrid infrastructure rather than just a massive GPU cluster?

**Dr. Robert Blumofe:** You know, this is a great point and you know, I think the more, I think people can really wrap their heads around what an agent really is architecturally, the better off we'll be. Because I think it's tempting to think that, well, you know, an AI agent is simply a super powerful LLM with the latest and greatest LLM reasoning capabilities. That's an agent, and that's not the case. The key insight probably is to think of an agent as being a system with many, many components. In fact, most agents do indeed have many, many components. And an LLM is just one of those components. The LLM. Typically, an LLM will typically play a fairly central role in the agent because you need something that's going to manage the natural language interaction and you need something that's going to make decisions about how the interaction should proceed, what's the next question to ask, what's the next task to do, and so on. So an LLM, or in many cases multiple LLMs play a fairly central role. But ultimately what makes these things agents is the ability to do things. And remember, an LLM can do nothing but produce text. If you want to do anything, you have to translate that text output into action. And that means tools. That's the key thing. LL, sorry, agents. Agents are systems that involve LLMs using tools. And in most good agents there's typically quite a few tools. It could be tools to read and write, storage. It could be tools to retrieve, retrieve information from say a, a CRM, right, to retrieve information about the customer that you're talking to. It may be a tool that retrieves information off of the web. It may be a tool that retrieves private information from a called vector database, could be a tool to send email, manage calendar. It could be a tool to calculate shortest paths on a map. So most agents are a combination of AI models, multiple AI models, plus a whole variety of tools. And I would argue, by the way, and I've been saying this for A while now that a good rule of thumb if you're designing an agent is to put as much of the functionality as possible into the non AI tools. You know, in some sense use AI only when nothing else will work, you know, and that doesn't mean don't use AI, of course, because the AI as the central component to making, making decisions and managing the natural language interaction. Well, nothing else will work. AI does that and it does it so well. But when it comes to other tasks, like things that I mentioned, like email, retrieving things from a database, searching the web. No, use it, use an actual tool, a non AI tool. It's way more efficient and way more reliable. So rule of thumb should be put as much of functionality in your agent as you can into the, into the non AI tools. Okay? The upshot of all that is that the infrastructure demands are coming from not just the LLM itself, but from the combination of multiple LLMs, multiple tools, using data, retrieving data. So you have a hybrid infrastructure requirement. You do need GPUs, but you also need CPU to run that shortest path algorithm to run the SQL query and so on. And, and of course you need, you need storage for all that data that you're going to be operating on, whether it's storing things like memories or retrieving things from a vector database. So you have this hybrid need. And by the way, you touched on this earlier, you know, another point I would make about sort of a good design role is for the parts that are AI, the parts that are, say, LLMs use the right LLM for the job. Not everything requires a multitrillion parameter ask me anything model. In many cases, if you're building a, an agent for a specific use, you really can, you're really going to be much better off with a, with an LLM that's much smaller and specialized for that particular, for that particular task. If you're building an agent to help your customers file insurance claims, you probably don't need an agent that can write code, compose sonnets, tell jokes, and give you the cast of every mash episode that ever was recorded. So, you know, use the right tool for the job and use the right AI for the job.

**Swapnil Bhartiya:** Akamai recently launched AI grid intelligent orchestration for distributed inference across your 4000 edge locations. What exactly is it and what does it mean in practice for developers?

**Dr. Robert Blumofe:** Yeah, thanks for the question. It's a great question. I would raise it to simply the basic notion of providing the right infrastructure in the right place at the right time. So starting with right infrastructure, it's not a one size fits all. It's not, as we said, it's not massive GPU cloud cluster for everything. That's good for some use cases and it's not just cpu. CPU again alone is good for some cases, but again, not for everything. And it's not just storage. It's the right combination of gpu, CPU and storage and connectivity for the use case. So that's the first question is you have to deliver the right infrastructure for the use case as it's presented to you at that time. Then there's the where. Again, it's not a one size fits all. You can't do everything in Ashburn, Virginia. It's deploying the right infrastructure in the right place for that particular use case. If the demand is coming from Dallas, Texas, infrastructure in Dallas, Texas, if it's using tools that are distributed in other locations, you want to have proximity to those tools. And then there's the connectivity. You need connectivity to all of those things. So it's the right infrastructure in the right place at the right time. There is no one size fits all for, for this stuff. And that's a challenge, by the way, because you know, it'd be nice if we could simply invest in a particular kind of infrastructure in a particular location. Problem solved. And it's just not going to work that way. It hasn't worked that way for the web. And that's certainly by the way. I think the cloud has done such a great job at this hybrid notion of infrastructure. I think that's one of the really great things about cloud is that it's not a one size fits all now. They're more centralized than we'd like them to be. But I think in terms of delivering the right type of infrastructure, I think that's one of the things that the cloud has really excelled at. You can choose what you're getting, the mix of CPU to GPU to, to storage so that you don't have to be stuck in that one size fits all. And I think that's the key, probably the key challenge, but the key recipe for success is recognizing that it's. That it's the, it's the right infrastructure, the right place at the right time. It's not a one size fits all. Not easy, but, but that's what needs to be delivered.

**Swapnil Bhartiya:** The current state of AI doesn't seem very sustainable. From the massive energy demands to token costs going through the roof. Looking two to three years ahead, how does AI infrastructure need to evolve to actually become sustainable?

**Dr. Robert Blumofe:** I really do think it comes down to sort of an intelligent sort of alternative approach to the, to the brute force approach. And the intelligent approach is actually fairly simple and it really is the things that we were just talking about, it's when you build your agent, it's use the right AI for the task. You don't have to do everything with an ask me anything multi trillion parameter model. It's using the right tools for the task, right? Use non AI whenever you can use non AI because that's much cheaper. It's delivering the right infrastructure to the task and delivering it in the right place. Doing all those things together can dramatically lower the cost and make these applications scalable and therefore much more usable. And I get the temptation to sort of, you know, use this brute force approach. And on the small scale maybe it's okay. You know, anecdotally, you know, I've been, I like to play with these agents and play with LLMs and I've been using things like Open Clon and Hermes Agent and I oftentimes find myself, I've got the thing configured to use Claude Opus 4.7 for example, which is a great, just a phenomenally great model. But then I'm sort of looking at the stuff that I'm doing with it, thinking wait a minute, you know, I don't need that level of model to do what I'm doing. So I'm racking up these ridiculous, you know, token fees and you know, okay, it's one thing for me to, you know, spend a little bit more money, you know, personally just for my own use, but if you tried to scale that to a real application that's going to be used by, by millions of people, using the wrong model is just a killer and using the wrong infrastructure is just a killer. So you've got to have the right models, the right tools, the right infrastructure in the right place. That intelligent approach is what makes the whole thing scale and is what ultimately is going to make AI ubiquitous. And it is going to be ubiquitous, you know, and we don't need any fancy new breakthroughs. We don't need AGI, we don't need quantum computing. AI as it lives today, with some good engineering and some good intelligent choices, can deliver some really phenomenal up levelings of the experience that we all have using computers or using any services online.

**Swapnil Bhartiya:** For enterprises already heavily invested in centralized cloud providers, what is the realistic path to this distributed model for them? Is it rip and replace or a more gradual transition?

**Dr. Robert Blumofe:** It's a great question. I actually think that we're still early enough and I don't think there's all that much lock in at this point. And there are some, for example, almost all the models support, for example the OpenAI API. So pretty much if your agent is the LLM part of your agent or the way that your agent interacts with the, the central LLM or other AI agents is using that interface, well then it's pretty easy to change, swap out model providers. So I don't know that lock in right now is a, is a big concern. It might be if we don't sort of change our path within the next couple of years, but I don't think it's a big concern right now. So I really, I really think right now it's about really understanding how to build agents and how to design great agentic experiences. And I've often said that there's no magic bullet here, there's no easy button here. Building a great system is still hard work. Even in the regime of Claude code, building a great system is still hard work that you've got to think through design, architecture, engineering. And I think if people simply recognize that and simply put in the effort to build a great agentic experience, it will be transformative.

**Swapnil Bhartiya:** As inference moves to the edge. Can you talk about what are some kind of new real world applications that will become possible that cannot be done today because of this centralized architecture?

**Dr. Robert Blumofe:** Yeah, I mean I actually think that every, every type of interaction that we do is a candidate for, for, for being agentic. Whether it's, you know, the way you use your desktop, you know, a simple example. Just the other, I told the story multiple times. Just the other day, you know, I came across an interesting website and I wanted to send the link to my wife and our youngest son. So I cut the, cut the link opened up the messaging, you know, typed in, you know, compose a new message, paste, send. Not that hardest thing in the world. But what's going through the back of my mind is why did I have to do all that? Why didn't I just say, hey, please send this link to my wife and youngest son? Done. Same thing as anytime I'm doing anything on the web, I'm sort of in the back of my mind having the same, wondering the same thing. You know, the other day I'm sort of looking at a car website and I was browsing through, you know, different car, car makes for this, this car website and I'm wondering like, why am I not just having a conversation with, with, with an AI expert that knows everything about These cars, all their configuration options, all their pluses and minuses, maybe if I've shopped there before, it knows about my preferences. And why isn't it, you know, able to show me, you know, what I'm interested in and we can have a conversation and it can show me the cars that I'm interested in, maybe with video, maybe customized for me and so on. Why am I still, you know, going through web pages and clicking on links? So I think pretty much everything that we do with our, with our desktops, everything that we do with our, on the web will be, be. Will turn into an agentic experience because it's just, it's doable, it's better. And, and, and I, I still to this day wonder why I don't have more of. I think it's coming very quickly, but it's clearly not yet arrived. I don't think it's that far. You know, I've often said that I think that the web has transformed the Internet and now we've got AI inference transforming the Web. And I think the transformation, you know, inference transforming the web will be every bit as profound, if not more so, as the web transform the Internet, you know, and most people, you know, certainly younger people, have no idea what the Internet was before the web. I don't think it's that much longer from now before we're explaining to young people what a web page was and what it was to click on a link. There's no reason for that anymore.

**Swapnil Bhartiya:** Tying it back to your CDN roots. Could we eventually see smaller models cache it locally at the edge, similar to how we cache web content? How does that traditional edge model translate to the AI space?

**Dr. Robert Blumofe:** Yeah, I think the edge can be used in a lot of different ways. In the context of an AI application, you could do some of the computation at the edge. For example, when we talk about an agent, as I said, it does many, many different things, not just invoking LLMs. Some of those things could be done at the edge at very, very low latency. Even some of the AI things you could be doing at the edge at low latency with relatively small models. Also that can include intelligent routing. You know, today, again, we're in a fairly static world in terms of what models we use, and we end up using the same model pretty much for every request that we make, every interaction. There's no reason for that. And a fairly simple model running at the edge could probably make some good decisions about where the request should be routed and you want to route for the right model. The right infrastructure in the right location. So all those things can be considered when you make a rapid routing decision at the edge. Obviously there's also security things that you do at very low latency at the edge. So again, ultimately, as you look at these AI applications, these AI agents, you break it down into many, many different components. And I think many of the components, probably not all of them, but many of the components I think actually will run at the edge and the ones that don't will get very, very low bandwidth, low latency, high bandwidth connectivity from the edge to whatever more centralized infrastructure you need for that particular use case. So it's a hybrid. You know, we oftentimes talk about the compute needs not being a one or the other, but, but as sort of a continuum, a hybrid. So there's some things that are in the core, some things that are at the edge and things in between, and ultimately they work together to create a low latency, high bandwidth, high quality user experience.

**Swapnil Bhartiya:** Robert, thank you so much. Just like the early Internet, we are waiting on the infrastructure to unlock the next massive wave of innovation and Akama is clearly leading that charge. Thank you for joining me and I look forward to chat with you again. Thank you, thank you.

**Dr. Robert Blumofe:** Thanks for the opportunity to share. These are topics that I really enjoy talking about, care about. So I do appreciate the opportunity to express.

공유