Background

为什么集中式 AI 无法实现大规模应用

Akamai 首席技术官 Robert Blumofe 博士阐述了为什么集中式基础架构无法满足无处不在的 AI 智能体的扩展需求。

要点汇总

推理应当部署在边缘

虽然集中式、高密度的 GPU 数据中心非常适合训练 AI 模型,但全球推理需要分布式边缘架构,以提供实际应用所需的低延迟和高带宽。

智能体是多组件系统

真正的 AI 智能体属于复合系统,其中 LLM 负责推理与决策,而专门的非 AI 工具(如数据库查询、搜索和电子邮件工具)负责执行实际工作。

优先考虑使用非 AI 工具和小型模型

为了实现可持续且具有成本效益的 AI 工程,开发人员应使用最精简、最专业且满足任务需求的模型,并尽可能地将智能体设计为使用更便宜、更可靠的非 AI 工具。

“智能体化”的转变将重塑 Web

正如 Web 曾经彻底改变了互联网一样,无处不在的 AI 推理将重新定义我们与技术的交互方式,用多模态、对话优先的智能体体验来取代静态的页面导航和链接。

带时间戳的摘要

[00:00–01:46] 集中式 AI 的局限性:主持人 Swapnil Bhartiya 开场指出了实时 AI 智能体过度依赖集中式集群所带来的问题。Robert Blumofe 博士解释道,尽管在模型训练的初始阶段,集中式 GPU 基础架构必不可少,但这种简单粗暴的集中式模式无法通过扩展来为“无处不在的 AI”的发展(即行业重心向推理、AI 应用以及主动式 AI 智能体的转移)提供支持。

[01:46–03:45] 与早期 Web 及网络安全发展史的相似之处:Blumofe 将当前的局面与 Web 发展初期(当时集中式托管导致了万维网变成“万维等”)以及网络安全技术的演进历程进行了类比。他阐述了 Akamai 当年如何运用数学、分布式系统和算法,从边缘侧交付应用来解决 Web 的可扩展性难题——而这一成功范式如今必须被应用到 AI 推理领域。

[03:45–05:20] 训练与推理基础架构的对比:Blumofe 基于训练与推理在物理层面上的“亲和性”,深入剖析了两者在架构需求上的差异。训练阶段需要靠近集中式数据集和高密度的 GPU 集群;推理阶段则需要尽可能靠近全球分布的最终用户,以支持动态的多模态交互。

[05:20–07:19] AI 智能体的混合架构:Blumofe 解释道,AI 智能体是由多个大语言模型以及非 AI 工具(例如存储、CRM 访问、向量数据库和 Web 搜索)组成的复杂系统。正因如此,智能体需要的是融合 GPU、CPU 和存储的混合基础架构,而不仅仅是庞大的纯 GPU 集群。

[07:19–08:34] Akamai AI Grid 智能编排:在谈及 Akamai 新近推出、覆盖 4,000 个边缘节点的 AI Grid 时,Blumofe 详细阐述了其核心目标:即在合适的时间和地点,提供最恰当的基础架构(GPU、CPU 与存储资源的定制化组合),而不是提供一种僵化且集中式的“一刀切”解决方案。

[08:34–10:04] 实现 AI 的可持续运营:为了应对不断飙升的 token 成本和能源需求,Blumofe 建议开发者放弃简单粗暴的算力堆砌模式。他提倡通过利用规模更小、高度专业化的模型来构建更聪明的智能体,并在处理基础操作任务时,优先使用效率更高、非 AI 类的常规工具。

[10:04–10:55] 云服务提供商的锁定与迁移路径:Blumofe 指出,由于整个行业仍处于发展初期,大规模的云服务提供商锁定现象尚未形成真正的阻碍。由于目前大多数模型都采用了标准的 OpenAI API,开发人员可以轻松切换提供商,并逐步过渡到分布式模型。

[10:55–12:00] 未来的实际智能体应用场景:Blumofe 描绘了日常数字体验将如何向智能体界面演进。用户将不再需要手动复制链接或在静态网页之间点来点去,而是以对话优先的方式与专家级的 AI 智能体进行交互,而这些智能体能够动态地为用户呈现定制化的视频、图像和工具。

[12:00–13:21] 边缘计算与本地缓存:Blumofe 探讨了传统的边缘计算模式将如何平移应用到 AI 领域。他提出,诸如运行小型模型、执行安全检查,以及为智能匹配最佳的后端计算资源而进行快速路由决策等任务,都可以直接在边缘端高效完成。

视频脚本

视频脚本

**Swapnil Bhartiya:** AI conversations usually revolve around massive GPUs and centralized clusters. But real world AI agents demand millisecond response time at the point of contact, and centralization simply cannot scale to meet that moment. Akamai is rethinking AI cloud infrastructure by building a distributed grid for inference across 4,000 global locations. Joining us today once again is Dr. Robert Blumofe, EVP and CTO at Akamai. Robert, it's great to have you on the show.

**Dr. Robert Blumofe:** Thanks. Thanks for having me.

**Swapnil Bhartiya:** We are seeing massive investment in centralized AI data center right now. Why do you believe this centralized everything approach is the wrong model for the future of AI?

**Dr. Robert Blumofe:** So it's a great question and ultimately the central thesis is simply that the sort of brute force approach of building out large amounts of infrastructure in centralized locations ultimately, well, it's expensive. But ultimately even that expense aside, can't achieve the scale that's going to be needed as AI sort of moves into its next phase that you might characterize as ubiquitous AI. And I think it's worth maybe highlighting a couple of ways in which the demand has changed, because ultimately you need to look at the demand and see how the infrastructure aligns to that demand. And I would focus maybe on two shifts. One would be the shift from training to inference, and the other one I would characterize as the shift from sort of the early days of a chatbot to an AI application or an AI agent. You know, it wasn't that long ago focusing on the first of those shifts. It wasn't that long ago that most of the infrastructure demand really came from the training use case where you were. And by and large I'm talking about the pre training of large generative models like LLMs that was driving a huge amount of the infrastructure demand. And in that use case, absolutely centralized, large scale, dense GPU infrastructure makes a whole lot of sense. But as you move into inference, it changes a lot. And of course, and I think we all know this, that training is really sort of a mandatory cost that is necessary to realize the value through inference. All the value in AI comes from the inference. And training is simply an investment that we have to make to realize the return that you get through through inference. And now as we're moving into a more mature phase, much more of the demand is coming from inference. And that's a good thing because again, that's where we get the value. So inference driving demand is a very, is a very good thing. And I would argue that the nature of inference is changing quite a bit. And again, that's the shift I'm talking about from the chatbot to the AI application or the AI agent. You know, in the case of the chatbot, I think we really thought of AI as sort of a destination. It was intentional. You went to chatgpt.com to use AI or you fired up your anthropic Claude desktop to use AI. It was intentional. It was a destination. Once you move into AI powered applications and AI agents, it becomes ubiquitous. It's no longer a specific destination, it's just part of everything that you do. Certainly everything that you do online, you go to a website to look for a car AI, you go to a healthcare provider to make an appointment to see your doctor. AI, everything that you're doing is AI powered, probably even everything that you're doing on your desktop, even irrespective of the web, you know, you want to send something to, to your kids AI. So AI becomes ubiquitous, that changes the nature of the demand. And in that world where AI is ubiquitous, being used all the time by everyone, a centralized approach just really isn't, isn't going to cut it. And we risk sort of revisiting the old. You know, back then we called it the world wide wait. It could turn into large language molasses for lack of a better term.

**Swapnil Bhartiya:** If AI training is more or less like writing software, inference is like deploying it globally so people can use requires totally different infrastructure. You have handled high latency sensitive use cases like live sports and global cybersecurity. How do the lessons from those challenges apply to AI infrastructure today?

**Dr. Robert Blumofe:** Yeah, that's a great point. And I do see a lot of parallels to the early days of the web. Also, I think some of the changes that happened maybe a decade or so ago in cybersecurity. So in many ways people say this all the time. History does repeat itself and it's kind of repeating itself for the umpteenth time here. So while there's obvious differences, there's a lot about what's happening now that I think does parallel what we saw in the early days of the web. As the web was getting popular, really transforming the Internet, there were a lot of concerns that the web simply wouldn't scale to meet the demand. And in many ways those concerns probably were well founded because you did have a situation where web applications were centralized. Now we were pre cloud, but we did have hosting providers. And arguably the hosting providers back then were even more centralized than today's hyperscalers. By and large, most of the infrastructure was in the U.S. for example, was heavily located in places like Ashburn and San Jose. So every time you used a web application, you had to traverse large distances into a handful of centralized locations. And while that might have been okay in the very early days, where a website was a pretty static thing, just some text, maybe a few images, as you move into video, for example, and large demand for that video, you simply cannot meet the bandwidth requirements and the latency requirements through that centralized model. And that's really, I think, what led people to, you know, jokingly say that the World wide Web should be, you know, called the worldwide, worldwide wait. And people speculated that the web would simply collapse. And really that concern was the beginning of Akamai, where, you know, Tom and Danny, the two founders, came forward with a better approach, math algorithms, distributed systems, rather than brute force. And they showed that you can actually deliver websites and web applications from the edge of the Internet, dramatically increasing the available bandwidth, dramatically lowering the latency. And really that's what made the web work. And that was a critical ingredient. Also, as the web transitioned from these static sites to dynamic, where your communication is happening all the time, it's not just click on a link and wait for a response. You're constantly interacting with these web applications. CDNs made all of that work. And a similar approach, really the same approach, is in many ways what enabled powerful cybersecurity defenses. Because cybersecurity also went through a pretty strong transformation about 10 years ago, maybe a bit less. Where you moved from our biggest concern being things like Anonymous to sophisticated ransomware and the world of sophisticated attackers. With, with ransomware and powerful DDoS, extortion attacks, the centralized approach just wouldn't work. And again, you have to borrow from this playbook of math, distributed systems, algorithms, and that worked. AI today, I think, is in a very similar regime where as you move from training to inference, as you move from fairly low bandwidth and high latency types of interactions like the chatbot, where you're simply typing some text, waiting for a response type, typing some text, waiting for a response, you move from that into an AI powered application or an AI agent, the nature of the demand just changes dramatically. It becomes ubiquitous. It's constant, it's high bandwidth, it requires low latency. And the, again, the brute force approach just isn't going to work. You can't do this with purely centralized infrastructure. The demand has moved to the edge. So the infrastructure and capabilities of AI have to also move to the edge.

**Swapnil Bhartiya:** Let's look at training versus inference. Most companies care far more about inference to actually deliver AI to their users. Do these two faces genuinely demand completely different infrastructure architectures?

**Dr. Robert Blumofe:** Yeah, it's a great point and a great way to distinguish these use cases. And I think it's helpful to think about what is the important affinity. What does the infrastructure need in terms of proximity and, and arguably in the case of training, the key affinity, the key proximity requirements is the data set, the training data set. And most training data sets are fairly large and they're generally stored in some large storage cluster that's going to be fairly centralized. You typically wouldn't have your training data set distributed around a large number of locations. It's going to be fairly centralized, so it makes sense to do the training where the training data set is. Um, it's also the case that if you look at the actual computation, you know, it's very GPU dense. So a dense GPU cluster centralized where the training data set is, that makes a whole lot of sense when you move into inference. Well, what's the affinity? What does it need to be near? Well, it needs to be near the things that it's interacting with. And, and there's a lot of things that, that AI applications, AI agents have to interact with. But obviously one of the important users, of course, is the people. Us. You know, we are going to use agents to get things done for us. We're going to engage in conversations with these agents to help specify what it is we want done, to look at results, to review results, provide feedback. It's going to be very conversational. So the affinity of an agent, I mean, we can get into this a little bit more in a little bit because there's a lot of different affinities, but one of them clearly is to the users. And users are typically distributed over a fairly large swath of geography, whether it's a country or a continent or, or the entire world. So it makes no sense really for the agent that you and I are interacting with to be thousands of miles away, centralized in a single location. And relative also to what I was mentioning earlier about the nature of the interaction changing with the web, where we went to high definition video and things like that. The same thing is the case with agents. You don't want to think of an agent interaction as being simply text or even voice. A good agent is going to show us video, is going to show us images, and it's going to be dynamically updating the video and dynamically updating the images. These are capabilities that you don't get outside of AI. And it's one of the great benefits of using an AI agent is that you have all these modalities available, video, images, interaction that isn't available in with other technologies. So you know, as we see agents become more ubiquitous, I think we'll see these high bandwidth forms of interaction really take, take hold because they really deliver value and they deliver something compelling and interesting. And there's just no way to do that from through a brute force build out in centralized infrastructure.

**Swapnil Bhartiya:** When we talk about AI agents, everyone immediately fixates on GPU scarcity. Can you explain why AI agents actually require a hybrid infrastructure rather than just a massive GPU cluster?

**Dr. Robert Blumofe:** You know, this is a great point and you know, I think the more, I think people can really wrap their heads around what an agent really is architecturally, the better off we'll be. Because I think it's tempting to think that, well, you know, an AI agent is simply a super powerful LLM with the latest and greatest LLM reasoning capabilities. That's an agent, and that's not the case. The key insight probably is to think of an agent as being a system with many, many components. In fact, most agents do indeed have many, many components. And an LLM is just one of those components. The LLM. Typically, an LLM will typically play a fairly central role in the agent because you need something that's going to manage the natural language interaction and you need something that's going to make decisions about how the interaction should proceed, what's the next question to ask, what's the next task to do, and so on. So an LLM, or in many cases multiple LLMs play a fairly central role. But ultimately what makes these things agents is the ability to do things. And remember, an LLM can do nothing but produce text. If you want to do anything, you have to translate that text output into action. And that means tools. That's the key thing. LL, sorry, agents. Agents are systems that involve LLMs using tools. And in most good agents there's typically quite a few tools. It could be tools to read and write, storage. It could be tools to retrieve, retrieve information from say a, a CRM, right, to retrieve information about the customer that you're talking to. It may be a tool that retrieves information off of the web. It may be a tool that retrieves private information from a called vector database, could be a tool to send email, manage calendar. It could be a tool to calculate shortest paths on a map. So most agents are a combination of AI models, multiple AI models, plus a whole variety of tools. And I would argue, by the way, and I've been saying this for A while now that a good rule of thumb if you're designing an agent is to put as much of the functionality as possible into the non AI tools. You know, in some sense use AI only when nothing else will work, you know, and that doesn't mean don't use AI, of course, because the AI as the central component to making, making decisions and managing the natural language interaction. Well, nothing else will work. AI does that and it does it so well. But when it comes to other tasks, like things that I mentioned, like email, retrieving things from a database, searching the web. No, use it, use an actual tool, a non AI tool. It's way more efficient and way more reliable. So rule of thumb should be put as much of functionality in your agent as you can into the, into the non AI tools. Okay? The upshot of all that is that the infrastructure demands are coming from not just the LLM itself, but from the combination of multiple LLMs, multiple tools, using data, retrieving data. So you have a hybrid infrastructure requirement. You do need GPUs, but you also need CPU to run that shortest path algorithm to run the SQL query and so on. And, and of course you need, you need storage for all that data that you're going to be operating on, whether it's storing things like memories or retrieving things from a vector database. So you have this hybrid need. And by the way, you touched on this earlier, you know, another point I would make about sort of a good design role is for the parts that are AI, the parts that are, say, LLMs use the right LLM for the job. Not everything requires a multitrillion parameter ask me anything model. In many cases, if you're building a, an agent for a specific use, you really can, you're really going to be much better off with a, with an LLM that's much smaller and specialized for that particular, for that particular task. If you're building an agent to help your customers file insurance claims, you probably don't need an agent that can write code, compose sonnets, tell jokes, and give you the cast of every mash episode that ever was recorded. So, you know, use the right tool for the job and use the right AI for the job.

**Swapnil Bhartiya:** Akamai recently launched AI grid intelligent orchestration for distributed inference across your 4000 edge locations. What exactly is it and what does it mean in practice for developers?

**Dr. Robert Blumofe:** Yeah, thanks for the question. It's a great question. I would raise it to simply the basic notion of providing the right infrastructure in the right place at the right time. So starting with right infrastructure, it's not a one size fits all. It's not, as we said, it's not massive GPU cloud cluster for everything. That's good for some use cases and it's not just cpu. CPU again alone is good for some cases, but again, not for everything. And it's not just storage. It's the right combination of gpu, CPU and storage and connectivity for the use case. So that's the first question is you have to deliver the right infrastructure for the use case as it's presented to you at that time. Then there's the where. Again, it's not a one size fits all. You can't do everything in Ashburn, Virginia. It's deploying the right infrastructure in the right place for that particular use case. If the demand is coming from Dallas, Texas, infrastructure in Dallas, Texas, if it's using tools that are distributed in other locations, you want to have proximity to those tools. And then there's the connectivity. You need connectivity to all of those things. So it's the right infrastructure in the right place at the right time. There is no one size fits all for, for this stuff. And that's a challenge, by the way, because you know, it'd be nice if we could simply invest in a particular kind of infrastructure in a particular location. Problem solved. And it's just not going to work that way. It hasn't worked that way for the web. And that's certainly by the way. I think the cloud has done such a great job at this hybrid notion of infrastructure. I think that's one of the really great things about cloud is that it's not a one size fits all now. They're more centralized than we'd like them to be. But I think in terms of delivering the right type of infrastructure, I think that's one of the things that the cloud has really excelled at. You can choose what you're getting, the mix of CPU to GPU to, to storage so that you don't have to be stuck in that one size fits all. And I think that's the key, probably the key challenge, but the key recipe for success is recognizing that it's. That it's the, it's the right infrastructure, the right place at the right time. It's not a one size fits all. Not easy, but, but that's what needs to be delivered.

**Swapnil Bhartiya:** The current state of AI doesn't seem very sustainable. From the massive energy demands to token costs going through the roof. Looking two to three years ahead, how does AI infrastructure need to evolve to actually become sustainable?

**Dr. Robert Blumofe:** I really do think it comes down to sort of an intelligent sort of alternative approach to the, to the brute force approach. And the intelligent approach is actually fairly simple and it really is the things that we were just talking about, it's when you build your agent, it's use the right AI for the task. You don't have to do everything with an ask me anything multi trillion parameter model. It's using the right tools for the task, right? Use non AI whenever you can use non AI because that's much cheaper. It's delivering the right infrastructure to the task and delivering it in the right place. Doing all those things together can dramatically lower the cost and make these applications scalable and therefore much more usable. And I get the temptation to sort of, you know, use this brute force approach. And on the small scale maybe it's okay. You know, anecdotally, you know, I've been, I like to play with these agents and play with LLMs and I've been using things like Open Clon and Hermes Agent and I oftentimes find myself, I've got the thing configured to use Claude Opus 4.7 for example, which is a great, just a phenomenally great model. But then I'm sort of looking at the stuff that I'm doing with it, thinking wait a minute, you know, I don't need that level of model to do what I'm doing. So I'm racking up these ridiculous, you know, token fees and you know, okay, it's one thing for me to, you know, spend a little bit more money, you know, personally just for my own use, but if you tried to scale that to a real application that's going to be used by, by millions of people, using the wrong model is just a killer and using the wrong infrastructure is just a killer. So you've got to have the right models, the right tools, the right infrastructure in the right place. That intelligent approach is what makes the whole thing scale and is what ultimately is going to make AI ubiquitous. And it is going to be ubiquitous, you know, and we don't need any fancy new breakthroughs. We don't need AGI, we don't need quantum computing. AI as it lives today, with some good engineering and some good intelligent choices, can deliver some really phenomenal up levelings of the experience that we all have using computers or using any services online.

**Swapnil Bhartiya:** For enterprises already heavily invested in centralized cloud providers, what is the realistic path to this distributed model for them? Is it rip and replace or a more gradual transition?

**Dr. Robert Blumofe:** It's a great question. I actually think that we're still early enough and I don't think there's all that much lock in at this point. And there are some, for example, almost all the models support, for example the OpenAI API. So pretty much if your agent is the LLM part of your agent or the way that your agent interacts with the, the central LLM or other AI agents is using that interface, well then it's pretty easy to change, swap out model providers. So I don't know that lock in right now is a, is a big concern. It might be if we don't sort of change our path within the next couple of years, but I don't think it's a big concern right now. So I really, I really think right now it's about really understanding how to build agents and how to design great agentic experiences. And I've often said that there's no magic bullet here, there's no easy button here. Building a great system is still hard work. Even in the regime of Claude code, building a great system is still hard work that you've got to think through design, architecture, engineering. And I think if people simply recognize that and simply put in the effort to build a great agentic experience, it will be transformative.

**Swapnil Bhartiya:** As inference moves to the edge. Can you talk about what are some kind of new real world applications that will become possible that cannot be done today because of this centralized architecture?

**Dr. Robert Blumofe:** Yeah, I mean I actually think that every, every type of interaction that we do is a candidate for, for, for being agentic. Whether it's, you know, the way you use your desktop, you know, a simple example. Just the other, I told the story multiple times. Just the other day, you know, I came across an interesting website and I wanted to send the link to my wife and our youngest son. So I cut the, cut the link opened up the messaging, you know, typed in, you know, compose a new message, paste, send. Not that hardest thing in the world. But what's going through the back of my mind is why did I have to do all that? Why didn't I just say, hey, please send this link to my wife and youngest son? Done. Same thing as anytime I'm doing anything on the web, I'm sort of in the back of my mind having the same, wondering the same thing. You know, the other day I'm sort of looking at a car website and I was browsing through, you know, different car, car makes for this, this car website and I'm wondering like, why am I not just having a conversation with, with, with an AI expert that knows everything about These cars, all their configuration options, all their pluses and minuses, maybe if I've shopped there before, it knows about my preferences. And why isn't it, you know, able to show me, you know, what I'm interested in and we can have a conversation and it can show me the cars that I'm interested in, maybe with video, maybe customized for me and so on. Why am I still, you know, going through web pages and clicking on links? So I think pretty much everything that we do with our, with our desktops, everything that we do with our, on the web will be, be. Will turn into an agentic experience because it's just, it's doable, it's better. And, and, and I, I still to this day wonder why I don't have more of. I think it's coming very quickly, but it's clearly not yet arrived. I don't think it's that far. You know, I've often said that I think that the web has transformed the Internet and now we've got AI inference transforming the Web. And I think the transformation, you know, inference transforming the web will be every bit as profound, if not more so, as the web transform the Internet, you know, and most people, you know, certainly younger people, have no idea what the Internet was before the web. I don't think it's that much longer from now before we're explaining to young people what a web page was and what it was to click on a link. There's no reason for that anymore.

**Swapnil Bhartiya:** Tying it back to your CDN roots. Could we eventually see smaller models cache it locally at the edge, similar to how we cache web content? How does that traditional edge model translate to the AI space?

**Dr. Robert Blumofe:** Yeah, I think the edge can be used in a lot of different ways. In the context of an AI application, you could do some of the computation at the edge. For example, when we talk about an agent, as I said, it does many, many different things, not just invoking LLMs. Some of those things could be done at the edge at very, very low latency. Even some of the AI things you could be doing at the edge at low latency with relatively small models. Also that can include intelligent routing. You know, today, again, we're in a fairly static world in terms of what models we use, and we end up using the same model pretty much for every request that we make, every interaction. There's no reason for that. And a fairly simple model running at the edge could probably make some good decisions about where the request should be routed and you want to route for the right model. The right infrastructure in the right location. So all those things can be considered when you make a rapid routing decision at the edge. Obviously there's also security things that you do at very low latency at the edge. So again, ultimately, as you look at these AI applications, these AI agents, you break it down into many, many different components. And I think many of the components, probably not all of them, but many of the components I think actually will run at the edge and the ones that don't will get very, very low bandwidth, low latency, high bandwidth connectivity from the edge to whatever more centralized infrastructure you need for that particular use case. So it's a hybrid. You know, we oftentimes talk about the compute needs not being a one or the other, but, but as sort of a continuum, a hybrid. So there's some things that are in the core, some things that are at the edge and things in between, and ultimately they work together to create a low latency, high bandwidth, high quality user experience.

**Swapnil Bhartiya:** Robert, thank you so much. Just like the early Internet, we are waiting on the infrastructure to unlock the next massive wave of innovation and Akama is clearly leading that charge. Thank you for joining me and I look forward to chat with you again. Thank you, thank you.

**Dr. Robert Blumofe:** Thanks for the opportunity to share. These are topics that I really enjoy talking about, care about. So I do appreciate the opportunity to express.

分享