KAIST and Meta Unveil Revolutionary AI Data Center Design Enhancing Efficiency and Scalability

September 9, 2026
KAIST and Meta Unveil Revolutionary AI Data Center Design Enhancing Efficiency and Scalability
  • KAIST and Meta, with Panmnesia, are proposing a next-generation AI data center architecture that treats the entire facility as a single, coordinated computing resource by interconnecting CPUs, AI accelerators, and memory across the data center using CXL, effectively turning the center into one giant chip.

  • The approach prioritizes coordinating inter- and intra-rack connections to minimize latency variation, rather than focusing solely on raw per-device performance or intra-rack link speeds.

  • Future work includes exploring optical interconnects to boost speed and scalability and moving toward commercial deployment.

  • The architecture enables dynamic resource utilization, allowing replacement of failed components and reusing idle compute and memory resources for other tasks.

  • Resources are organized hierarchically into trays, pods, and fabric to maintain fixed-hop, consistent communication paths and timing across the data center.

  • Compared with Nvidia’s GB200 NVL72 approach, the Panmnesia/Meta concept claims up to an eightfold increase in coordinated accelerators per CPU and as many as 960 accelerators operating within a single coherence domain, with data access latencies in the hundreds of nanoseconds.

  • The architecture envisions connecting up to 960 accelerators within one coherence domain, roughly 13 times larger than a typical NVLink-based rack, enabling broad resource sharing.

  • The emphasis shifts from individual GPU speed to improvements in interconnect performance and data transfer efficiency as the main determinants of data center performance for large AI models.

  • Data-movement latency is expected to drop to hundreds of nanoseconds, improving overall computing efficiency as the system scales.

  • The work responds to the growing need for fast, scalable interconnection as AI models require thousands of accelerators for a single computation.

  • Fabric components have been validated in silicon, including a fabric controller and LAU, with a silicon fabric switch ready in pre-release, signaling manufacturability of the design.

  • By grouping many accelerators within a single connectivity domain, the design enables joint utilization across CPUs, accelerators, and memory, expanding data-center scale well beyond current NVLink racks.

Summary based on 3 sources


Get a daily email with more AI stories

More Stories