AI workloads are outgrowing the standalone data centers. Discover how a “scale-across” Ethernet fabric unifies distributed clusters, backed by Cisco Zeus AI Lab benchmarking, to ensure peak performance and ROI across any distance.
For years, organizations have leveraged AI/ML for specific tasks—from computer vision in video analytics to Google’s BERT powering advanced search and ML models driving predictive analytics. Industry leaders have long pioneered research in natural language processing (NLP) models for speech recognition, text summarization, and sentiment analysis, but the launch of ChatGPT marked the beginning of the AI renaissance and the generative AI era. Overnight, AI shifted from a futuristic concept to an essential daily collaborative tool. The rise of large language models (LLMs) and generative AI has further accelerated global innovation, disrupted established workflows, and forced every major industry to reimagine products and solutions.
Today, every organization—from small enterprises to the largest cloud providers—must decide how to integrate this technology and stay ahead amidst the rapidly evolving AI landscape.
The foundations of AI fabricsThe building blocks of a high-performance AI cluster environment include accelerators, storage servers, and the network fabrics that connect these servers (Figure 1).
Figure 1. AI fabric architecture for north-south and east-west connectivity
The frontend fabric: This acts as the gateway to your GPU cluster, handling standard data center traffic, including north-south user access, API calls, logging, and data ingestion from storage or data lakes into the GPUs.
The backend fabric: Also known as the “scale-out” fabric, this provides a high-speed, low-latency interconnect for GPU-to-GPU collective communication (east-west traffic).
As AI infrastructure requirements grow, hyperscalers, neoclouds, sovereign clouds, and large enterprises must rethink how they build AI clusters. Faced with limited power and space at individual locations, operators must build distributed data centers to meet the surging training and inference demands of trillion-parameter LLMs.
Because a single cluster cannot handle workloads of this scale, a distributed ecosystem supported by a scale-across fabric is required.
Unifying the distributed ecosystem with scale-across fabricAs data centers supporting AI/ML applications become more distributed, the scale-across fabric serves as a vital interconnect. This architecture allows AI clusters to transcend single-site limits, spanning multiple facilities to overcome space and power constraints. By extending geographically, operators can optimize cost and energy while unlocking capacity pooling, extended resilience, and seamless, continuous growth. By interconnecting data centers spanning hundreds of kilometers, operators aren’t just adding links—they are forming a mega-scale AI infrastructure that spans multiple locations.
Whatever fabric architecture is in place, Cisco champions an Ethernet-based approach because it provides the scale, interoperability, and ecosystem maturity required for modern AI infrastructure (Figure 2). As interface speeds evolve to 1.6 Tbps and beyond, Ethernet remains the only viable path that offers the massive east-west bandwidth and performance needed to future-proof AI investment.
Figure 2. Ethernet unifies the AI operating model—from scale-out to scale-across
Designing a scale-across fabric for optimal performance
Scale-across is not a one-size-fits-all architectural solution. Every environment is unique, and design constraints from distance to traffic patterns can significantly impact AI cluster performance. To architect a scale-across architecture effectively, several critical questions must first be addressed, including:
The core purpose of a scale-across fabric is to connect multiple AI clusters, so they behave as a single logical entity (Figure 3). However, because the frontend and backend fabrics serve completely different purposes, extending them introduces distinct architectural challenges.
Figure 3. Connecting frontend and backend AI/ML fabrics across data centers
Understanding the variables of distributed AI success
When network reach is extended beyond one data center, the underlying design principles change. Organizations are no longer dealing with one simple network topology. Four critical variables define the success of this architecture, including:
Interconnecting distributed data center locations introduces profound architectural complexities (Figure 4). Enterprises need not face interconnect hurdles through trial and error.
Figure 4: Cisco provides a proven blueprint to keep distributed AI workloads stable and optimized at any scale
Cisco’s dedicated AI Infrastructure Benchmarking Engineering Team conducts rigorous, empirical research to establish proven, validated reference designs. This systematic testing methodology encompasses critical networking features, including:
Performance evaluation tools and key metrics
Test topology and architectural use cases
Physical and operational environments
Performance tuning and infrastructure optimization
Although formal industry standards for distributed AI fabrics are not yet established, research defines three distinct tiers to guide AI infrastructure architecture (Figure 5):
Figure 5: How distance changes the scale-across architecture design
The future of AI infrastructure is no longer defined solely by isolated computational capacity; it is defined by fabric connectivity that transcends physical geographical boundaries. As organizations refine data center strategies for the next wave of innovation, the defining architectural question remains: Is your scale-across strategy prepared for what comes next?
The evolution of mega-scale AI infrastructure introduces unprecedented complexities. In an upcoming technical blog series, we will deconstruct each of these scale-across architectural tiers. This series will deliver empirical insights and data-driven recommendations derived from the Cisco Zeus AI Lab, enabling organizations to deploy and scale distributed clusters with predictability and confidence.
Power your network for AI at scale with Cisco AI Networking
Additional resources:
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | Scaling the future: Why Ethernet is the backbone of AI Supercomputing | 0 | 9.17 | 06-08-2026 |
| 2 | Is your SD-WAN ready for AI-powered operations? | 0 | 18.01 | 03-08-2026 |
| 3 | As Goes AI Compute, So Goes Ethernet Networking | 0 | 7.66 | 06-07-2026 |
| 4 | Protecting against rising cybersecurity risks in data centers | 0 | 13.33 | 30-06-2026 |
| 5 | Cisco Nexus One, next-generation data center networking architecture | 0 | 10.4 | 02-07-2026 |
| 6 | From AI Experiments to 90% Adoption: How Cisco Operationalized AI at Scale | 0 | 14.13 | 27-07-2026 |
| 7 | Unifying operations with Cisco Cloud Control | 0 | 10.97 | 09-07-2026 |
| 8 | Why AI Infrastructure Is The Key To Enterprise AI Success | 0 | 6.21 | 21-04-2026 |
| 9 | One fallen power line exposed a growing AI data center problem — here’s how to fix it | 0 | 5.82 | 25-07-2026 |
| 10 | Arista Rides AI Scale Out Networks, Moves Into Scale Across, And Awaits Scale Up | 0 | 7.35 | 07-05-2026 |