AMD has introduced its new GPU architecture, CDNA 5, designed specifically for servers. This architecture not only builds upon its predecessors but also introduces some surprises. For instance, parts of its internal configuration align closer with RDNA, concurrently abandoning innovations like the Infinity Cache.
CDNA 5: A Distinct Evolution
CDNA 5 represents more than a mere incremental upgrade from previous CDNA architectures. Since the introduction of CDNA as a successor to GCN, its focus on compute performance has set it apart from RDNA, which is tailored for gaming optimizations. While GCN was originally designed for high computational loads in server and research environments, CDNA maintains this focus.
The Evolution of AMD’s Compute Architecture
CDNA has retained the use of HBM for graphics memory and adopted AI cores similar to Nvidia’s tensor cores, optimizing throughput and computational power. Conversely, RDNA has been fine-tuned for gaming performance and efficient resource allocation.
As AI developed into a powerhouse in 2024, CDNA 3 expanded its data format offerings to keep pace with the training of Large Language Models (LLMs). Earlier, only FP16 was supported, but with CDNA 3, INT8 and FP8 formats were accelerated for matrix computations. Alongside, the sparsity function—popularized by Nvidia—was included to further double performance in certain tasks. CDNA 3 also integrated the successful Infinity Cache, which buffers access to graphics memory similar to CPU operations.
Smaller, Faster, and More Efficient
In CDNA 5, AMD is making significant advancements. All chiplets have undergone a shrink process, while the architecture has improved significantly, speeding up memory access as well. AMD has published a comprehensive white paper detailing these changes, following extensive coverage from AMD’s Advancing AI event.
Introducing Work Group Processors
The chiplet housing the GPU cores is the XCD, or Accelerator Complex Die. Each chiplet features 32 computation cores, now renamed Work Group Processors (WGPs). This renaming is reminiscent of RDNA, where two CUs were combined into one WGP for improved task efficiency. Though AMD hasn’t disclosed specifics, the nearly doubled transistor count hints that the new WGPs likely offer double the throughput compared to the previous CUs.
Revamped Cache Structure
To connect the XCDs, AMD has switched to a dual Fabric and Cache Die (FCD) configuration. In a major change, L2 Cache has been centralized within the FCDs, eliminating the previous Infinity Cache concept and increasing the L2 cache size to 192 MB per GPU. With a speed of 27 TB/s, this setup ensures lower latencies crucial for AI workloads.
New Data Types and Instructions
AMD continues to evolve its data types in response to AI advancements. Newly introduced micro-scale data types (MXFP4-MXFP8) enhance efficiency by allowing shared exponents among numbers, ultimately saving memory. Additionally, AMD’s support for the hyperbolic tangent function (tanh) in hardware accelerates machine learning processes, promising a significant performance boost.
The Competition with Nvidia
AMD has outpaced Nvidia in several server GPU segments. Leveraging advanced manufacturing processes, it boasts larger GPUs and faster HBM. While Nvidia’s CUDA technology has long dominated, AMD’s ROCm have made strides, indicating that many clients no longer see CUDA as critical in their workflows. As AMD seeks to convert technological advantages into market share, the question remains: how long will Nvidia wait to unveil new products?
Found this article interesting or helpful? Your support is appreciated through ComputerBase Pro and by disabling ad blockers. Learn more about advertisements on ComputerBase.

