Synopsys Rolls Out First Complete CXL 4.0 IP for AI Memory Connectivity
The controller, PHY, security, and verification IP support 128 GT/s links and target rack-scale memory pools exceeding 100 TB.
Synopsys announced what it calls the industry's first complete CXL 4.0 IP solution, combining a controller, security modules, a silicon-proven PHY, and verification IP for the latest generation of the Compute Express Link interconnect. The release comes roughly nine months after the CXL Consortium published the 4.0 specification in November, and it targets the memory bandwidth and capacity limits that increasingly govern AI system performance.

The evolution of CXL. Image used courtesy of the Compute Express Link Consortium
CXL 4.0 doubles the link rate over CXL 3.x, from 64 GT/s to 128 GT/s, by adopting a PCIe 7.0-based physical layer while retaining the 256-byte FLIT format and full backward compatibility. The specification also adds support for up to four retimers per link and native x2 link widths, both aimed at extending fabric reach across a rack. Synopsys' package is designed to carry designs from that specification into working silicon across CXL 4.0, 3.x, 2.0, and 1.x on a single unified architecture.
Inside the IP Package
The controller implements the full CXL 4.0 feature set at 128 GT/s, including bundled ports, port-based routing, and the 256-byte latency-optimized FLIT. Bundled ports aggregate multiple links into one logical connection; Synopsys says bundling four x16 links yields more than 2 TB/s of bandwidth, and eight x16 links more than 4 TB/s. One license covers every CXL generation and includes PCIe 7.0 fallback without a separate PCIe license.

Synopsys' CXL IP solution. Image used courtesy of Synopsys
The PHY is a silicon-proven, PCIe 7.0-based SerDes running at 128 GT/s, hardened for process nodes from 5 nm to 2 nm, with low-jitter tolerance and low-latency forward error correction. The verification IP, per the company, is the first commercial VIP for CXL 4.0, aimed at keeping compliance and interoperability testing off the critical path. Synopsys grounds the offering in its interconnect track record: more than 3,800 PCIe design wins and over 170 CXL controllers and PHYs shipped.
Security Without a Latency Tax
The security piece addresses a specific tension in multi-tenant AI and cloud deployments. Operators need coherent memory traffic encrypted and authenticated, but CXL's value rests on latency. Synopsys quotes CXL.mem load-to-use latency below 200 ns, close to cache-coherent territory, and states that the jump to 128 GT/s adds no latency over CXL 3.x. Encryption logic that added cycles on those paths would consume the margin that the interconnect was designed to provide.
Synopsys' Integrity and Data Encryption (IDE) modules perform AES-GCM encryption and authentication with zero-cycle latency overhead on CXL.cache and CXL.mem traffic in skid mode, so protected links run at the same latency as unprotected ones.

CXL 4.0 doubles the bandwidth of its predecessors without increasing the latency. Image used courtesy of the Compute Express Link Consortium
The modules also support the TSP and TDISP protocols for confidential computing in virtualized environments and are described as ready for FIPS 140-3 certification, a common requirement in cloud and government procurement. Synopsys offers dedicated IDE modules for the 4.0, 3.x, and 2.0 generations of the specification, matching the controller's multi-generation coverage.
Attacking the AI Memory Wall
Synopsys is keen to associate the release with a shift in where AI systems bottleneck: LLM inference is dominated by memory access rather than raw compute. The company cites Goldman Sachs Research projections that global AI token consumption will grow 24 times by 2030, reaching roughly 120 quadrillion tokens per month, with the KV caches behind those tokens straining GPU-attached memory.
CXL 4.0 addresses the problem through disaggregation. Rack-scale memory pooling can exceed 100 TB of coherent memory, allowing servers to draw capacity on demand rather than overprovisioning every node. Synopsys claims KV cache offload over 128 GT/s CXL runs three to six times faster than SSD-based alternatives, and estimates that pooling cuts inference costs substantially.
The CXL positioning within Synopsys' broader HPC portfolio places it alongside PCIe 7.0, UCIe, UALink, and Ultra Ethernet IP, covering the scale-up and scale-out fabrics that AI accelerator designs now mix and match. Where earlier CXL generations established memory pooling (2.0) and switching fabrics (3.x), the company argues 4.0 is the generation where the interconnect's bandwidth finally matches the demands of disaggregated AI systems.
The move to CXL 4.0 is quite interesting, especially with the growing memory demands of AI workloads. Doubling the link rate to 128 GT/s while maintaining low latency could make memory pooling and disaggregation much more practical for large-scale AI systems. The combination of security, PHY, controller, and verification IP in a single solution also looks useful for developers working on next-generation accelerator and server designs.