NVIDIA GeForce RTX 5090 D

NVIDIA GeForce RTX 5090 D
NVIDIA GeForce RTX 5090 D graphics card review

NVIDIA GeForce RTX 5090 D: AI Cut Back, but Gaming Remains Intact

The GeForce RTX 5090 D has been launched for the Chinese market due to U.S. export restrictions. Unlike the RTX 4090 D, NVIDIA has hardly cut back on its gaming capabilities; the number of CUDA cores, memory, and core GPU units remain at the level of the standard RTX 5090. The main limitation concerns AI performance.

What Exactly Was Cut in the RTX 5090 D

The RTX 5090 D retains 21,760 CUDA cores, 680 Tensor Cores, 170 RT Cores, 32 GB of GDDR7 memory, a 512-bit memory bus, and a 575-watt power limit.

The main difference is the declared performance of the Tensor cores: 2375 AI TOPS for the RTX 5090 D compared to 3352 AI TOPS for the standard RTX 5090, which is approximately 29% lower.

In gaming, this limitation hardly affects performance: rasterization, ray tracing, and the memory subsystem of the RTX 5090 D remain intact.

Performance Details of the RTX 5090 D

The performance of partner models for the RTX 5090 D is particularly indicative.

Graphics Card 3DMark Time Spy Extreme Graphics 3DMark Port Royal
Colorful iGame GeForce RTX 5090 D Advanced 32GB 25,497 36,497
ZOTAC GeForce RTX 5090 D SOLID OC 32GB 25,901 35,252
GALAX GeForce RTX 5090 D Xingyao LUNA OC 24,946 34,017

The results were obtained on different test systems, so it's not advisable to compare models based on a few hundred points directly. The difference among the three cards is only a few percent-typical variability among different versions of the same GPU.

Factory overclocking offers only minor improvements, so when choosing a specific RTX 5090 D, it's better to focus on cooling, noise, dimensions, and power limits.

RTX 5090 D in 4K

4K is the most indicative mode for the RTX 5090 D: at 1440p and especially at 1080p, results are increasingly limited by the CPU or game engine.

Results with DLSS 4 Multi Frame Generation should be considered separately from regular tests: frame generation significantly boosts the FPS count, but the graphics card doesn't draw frames that much faster. Therefore, for direct comparisons, the RTX 5090 D's performance is better reflected in tests without Frame Generation.

Why the Original RTX 5090 D Needs 32 GB

The original RTX 5090 D received 32 GB of GDDR7, a 512-bit bus, and a bandwidth of 1792 GB/s.

In most modern games, 32 GB isn't always necessary, but such capacity is beneficial for 3D rendering, video editing, large scenes, and running local AI models.

However, for computational tasks, the RTX 5090 D is no longer a complete analog of the standard RTX 5090. The declared AI performance of the D version is lower, and the real difference depends on the application and format of computations.

In AI tasks, the limitations of the D version are far more significant than in gaming.

575 Watts Require Serious Power and Cooling

The RTX 5090 D retains the 575-watt TGP of the standard RTX 5090. Partner versions with expanded power limits can consume even more.

A powerful power supply, a well-ventilated case, and enough space for the massive graphics card will be necessary. Larger partner models occupy up to four slots and weigh more than 2.5 kg.

For such heavy cards, support for the graphics card is no longer just a decorative accessory.

RTX 5090 D and RTX 5090 D V2 - Not the Same Thing

The original RTX 5090 D comes with:

  • 32 GB of GDDR7;
  • 512-bit memory bus;
  • bandwidth of 1792 GB/s.

Later, the RTX 5090 D V2 was introduced: it has the same 21,760 CUDA cores but comes with only 24 GB of GDDR7 and a 384-bit bus. The memory bandwidth has dropped to 1344 GB/s.

The V2 was released after another tightening of export restrictions and effectively replaced the original version in the Chinese market.

In gaming, reduced memory and bandwidth typically have a lesser impact than in work workloads, but the difference is more noticeable for AI and heavy professional tasks. Therefore, the designation V2 should not be taken as a simple revision of the same card.

Conclusion

The original RTX 5090 D hardly falls short of the RTX 5090 in gaming capabilities: NVIDIA has transferred the main limitations to AI computations.

If the standard RTX 5090 is priced the same, there’s little point in choosing the D version. However, at a lower price, the original 32GB RTX 5090 D loses almost nothing in gaming.

The key when purchasing is to not confuse it with the RTX 5090 D V2. The V2 already has 24 GB of memory and a 384-bit bus: this is not a cosmetic revision, but a significantly altered version of the card.

Basic

Label Name
NVIDIA
Platform
Desktop
Launch Date
January 2025
Model Name
GeForce RTX 5090 D
Generation
GeForce 50
Base Clock
2017 MHz
Boost Clock
2407 MHz
Bus Interface
PCIe 5.0 x16
Transistors
92 billion
RT Cores
170
Tensor Cores
?
Tensor Cores are specialized processing units designed specifically for deep learning, providing higher training and inference performance compared to FP32 training. They enable rapid computations in areas such as computer vision, natural language processing, speech recognition, text-to-speech conversion, and personalized recommendations. The two most notable applications of Tensor Cores are DLSS (Deep Learning Super Sampling) and AI Denoiser for noise reduction.
680
TMUs
?
Texture Mapping Units (TMUs) serve as components of the GPU, which are capable of rotating, scaling, and distorting binary images, and then placing them as textures onto any plane of a given 3D model. This process is called texture mapping.
680
Foundry
TSMC
Process Size
4 nm
Architecture
Blackwell

Memory Specifications

Memory Size
32GB
Memory Type
GDDR7
Memory Bus
?
The memory bus width refers to the number of bits of data that the video memory can transfer within a single clock cycle. The larger the bus width, the greater the amount of data that can be transmitted instantaneously, making it one of the crucial parameters of video memory. The memory bandwidth is calculated as: Memory Bandwidth = Memory Frequency x Memory Bus Width / 8. Therefore, when the memory frequencies are similar, the memory bus width will determine the size of the memory bandwidth.
512bit
Memory Clock
1750 MHz
Bandwidth
?
Memory bandwidth refers to the data transfer rate between the graphics chip and the video memory. It is measured in bytes per second, and the formula to calculate it is: memory bandwidth = working frequency × memory bus width / 8 bits.
1.79TB/s

Display and Media

Outputs
1x HDMI 2.1b
3x DisplayPort 2.1b

Theoretical Performance

Pixel Rate
?
Pixel fill rate refers to the number of pixels a graphics processing unit (GPU) can render per second, measured in MPixels/s (million pixels per second) or GPixels/s (billion pixels per second). It is the most commonly used metric to evaluate the pixel processing performance of a graphics card.
423.6 GPixel/s
Texture Rate
?
Texture fill rate refers to the number of texture map elements (texels) that a GPU can map to pixels in a single second.
1637 GTexel/s
FP16 (half)
?
An important metric for measuring GPU performance is floating-point computing capability. Half-precision floating-point numbers (16-bit) are used for applications like machine learning, where lower precision is acceptable. Single-precision floating-point numbers (32-bit) are used for common multimedia and graphics processing tasks, while double-precision floating-point numbers (64-bit) are required for scientific computing that demands a wide numeric range and high accuracy.
104.8 TFLOPS
FP64 (double)
?
An important metric for measuring GPU performance is floating-point computing capability. Double-precision floating-point numbers (64-bit) are required for scientific computing that demands a wide numeric range and high accuracy, while single-precision floating-point numbers (32-bit) are used for common multimedia and graphics processing tasks. Half-precision floating-point numbers (16-bit) are used for applications like machine learning, where lower precision is acceptable.
1.637 TFLOPS
FP32 (float)
?
An important metric for measuring GPU performance is floating-point computing capability. Single-precision floating-point numbers (32-bit) are used for common multimedia and graphics processing tasks, while double-precision floating-point numbers (64-bit) are required for scientific computing that demands a wide numeric range and high accuracy. Half-precision floating-point numbers (16-bit) are used for applications like machine learning, where lower precision is acceptable.
104.8 TFLOPS

Miscellaneous

SM Count
?
Multiple Streaming Processors (SPs), along with other resources, form a Streaming Multiprocessor (SM), which is also referred to as a GPU's major core. These additional resources include components such as warp schedulers, registers, and shared memory. The SM can be considered the heart of the GPU, similar to a CPU core, with registers and shared memory being scarce resources within the SM.
170
Shading Units
?
The most fundamental processing unit is the Streaming Processor (SP), where specific instructions and tasks are executed. GPUs perform parallel computing, which means multiple SPs work simultaneously to process tasks.
21760
L1 Cache
128 KB (per SM)
L2 Cache
96 MB
TDP
575W
Vulkan Version
?
Vulkan is a cross-platform graphics and compute API by Khronos Group, offering high performance and low CPU overhead. It lets developers control the GPU directly, reduces rendering overhead, and supports multi-threading and multi-core processors.
1.4
OpenCL Version
3.0
OpenGL
4.6
CUDA
12.0
DirectX
12 Ultimate (12_2)
Power Connectors
1x 16-pin
ROPs
?
The Raster Operations Pipeline (ROPs) is primarily responsible for handling lighting and reflection calculations in games, as well as managing effects like anti-aliasing (AA), high resolution, smoke, and fire. The more demanding the anti-aliasing and lighting effects in a game, the higher the performance requirements for the ROPs; otherwise, it may result in a sharp drop in frame rate.
176
Shader Model
6.8
Suggested PSU
1000 W

Benchmarks

FP32 (float)
Score
104.8 TFLOPS
3DMark Steel Nomad
Score
14475
Blender
Score
14853.7
Vulkan
Score
382809
OpenCL
Score
385013

Compared to Other GPU

FP32 (float) / TFLOPS
163.4 +55.9%
104.8 +0%
79.478 -24.2%
65.572 -37.4%
3DMark Steel Nomad
14544 +0.5%
9237 -36.2%
8858 -38.8%
Blender
15026.3 +1.2%
1265.43 -91.5%
630 -95.8%
Vulkan
152166 -60.3%
100987 -73.6%
76392 -80%
49804 -87%
OpenCL
388405 +0.9%
126692 -67.1%
90580 -76.5%
66428 -82.7%