NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090
NVIDIA GeForce RTX 5090 graphics card review

NVIDIA GeForce RTX 5090: 32 GB GDDR7 at 575 W

The GeForce RTX 5090 is a graphics card that makes no pretense of being a rational purchase. Its tasks include 4K gaming at maximum settings, demanding ray tracing, local AI, and professional rendering. However, against the backdrop of a significant configuration update, the increase in standard gaming FPS has been relatively modest: the number of CUDA cores has increased by about a third, memory bandwidth has grown by nearly 78%, TGP has risen to 575 W, and the advantage over the RTX 4090 without frame generation typically stands at around 25-30%.

Therefore, traditional rendering and results with DLSS should be considered separately. Multi Frame Generation demonstrates the capabilities of Blackwell and DLSS but does not represent a twofold increase in GPU performance itself.

Specifications of the GeForce RTX 5090

At the heart of the RTX 5090 is Blackwell GB202. The card features 21,760 CUDA cores, 32 GB of GDDR7 memory, and a 512-bit memory bus. The bandwidth reaches 1792 GB/s-almost 1.8 TB/s. It also utilizes fifth-generation Tensor Cores and fourth-generation RT Cores.

Specification GeForce RTX 5090 GeForce RTX 4090
Architecture Blackwell Ada Lovelace
CUDA Cores 21,760 16,384
Video Memory 32 GB GDDR7 24 GB GDDR6X
Memory Bus 512 bits 384 bits
Bandwidth 1792 GB/s about 1008 GB/s
Boost Clock 2.41 GHz 2.52 GHz
PCI Express PCIe 5.0 PCIe 4.0
TGP 575 W 450 W
Starting Price $1999 $1599

The RTX 4090 claims a higher stated Boost clock, so comparing generations solely by gigahertz is meaningless. The RTX 5090 has about a third more CUDA cores, 8 GB more memory, and almost 78% higher bandwidth.

Benchmarks of the GeForce RTX 5090

Graphics Card 3DMark Time Spy Extreme Graphics To Founders Edition
NVIDIA GeForce RTX 5090 Founders Edition ~25,431 -
MSI GeForce RTX 5090 SUPRIM SOC ~26,345 +3.6%
ASUS ROG Astral GeForce RTX 5090 OC ~26,504 +4.2%
GIGABYTE AORUS GeForce RTX 5090 MASTER ~26,741 +5.2%

Even the most overclocked factory versions yield only a few percentage points advantage over the Founders Edition. In ray tracing tests, the gap remains similarly small.

Choosing between ROG Astral, SUPRIM, AORUS MASTER, and Founders Edition should be based on cooling, noise levels, size, and power limits rather than an expectation of significantly higher FPS. A massive cooling system does not make the GB202 significantly faster.

RTX 5090 vs. RTX 4090 in Games

At 4K, the RTX 5090 typically outperforms the RTX 4090 by about 25-30% without frame generation. In particularly demanding projects, the gap exceeds 30%, while in 1440p, it is often limited by the CPU.

In Black Myth: Wukong, at 4K, it achieves about 62 FPS compared to approximately 47 FPS for the RTX 4090-roughly a +32% advantage for the RTX 5090.

There is no doubling of RTX 4090's performance in standard rendering here.

The RTX 5090 is primarily needed for 4K gaming. At 1080p and often at 1440p, the CPU already limits the gains. The higher the resolution and the more demanding the ray tracing or path tracing, the better the computational power of the new GPU is utilized.

DLSS and Multi Frame Generation

Blackwell introduced Multi Frame Generation: multiple additional frames can be created between fully rendered frames. With this technology, the FPS counter can increase significantly.

However, one cannot equate such results with regular FPS.

Multi Frame Generation enhances visual smoothness but does not convert 60 original FPS into 200 FPS with corresponding control responsiveness. Therefore, results with MFG cannot be used as an indicator of pure GPU performance increase.

The maximum advantage of the RTX 5090 appears at 4K with heavy ray tracing, path tracing, and neural network rendering.

32 GB is Needed More for Work Tasks than for Gaming

For games, the additional 8 GB rarely changes the final FPS yet. In local AI and heavy 3D scenes, the difference between 24 and 32 GB can determine whether a task fits in video memory.

In Blender, the RTX 5090 is capable of outperforming the RTX 4090 by about 35-40%, and in some AI workloads, the increase is also noticeably higher than in standard rasterization.

The Blackwell architecture has support for FP4, and the claimed AI performance of the RTX 5090 reaches 3352 TOPS. In tandem with 32 GB of VRAM, this expands the range of models and workloads that can be fully run on the GPU.

The RTX 5090 also features three ninth-generation NVENCs and two sixth-generation NVDECs. Three hardware encoders are particularly beneficial for editing, streaming, and parallel encoding of multiple video streams.

575 W is the Main Compromise

The RTX 4090 could no longer be considered economical, but the RTX 5090 raised TGP from 450 to 575 W.

TGP has increased by about 28%, nearly in proportion to the average gaming performance increase at 4K. From a performance-per-watt perspective, this transition appears much more modest compared to the increase in the number of CUDA cores and memory bandwidth.

A power supply of about 1000 W is recommended for systems with the RTX 5090, although actual need depends on the CPU and other components.

Concurrently, the Founders Edition remains a dual-slot card about 304 mm long, thanks to a redesigned PCB and Double Flow Through cooling system. Most partner RTX 5090s are significantly larger.

Against the backdrop of massive three- and four-slot models, this is a significant advantage for the Founders Edition.

Early RTX 5090s Should Check ROP Count

Among early RTX 5090s, there were instances with fewer ROPs than specified.

A normal RTX 5090 should have 176 ROPs. According to NVIDIA, the issue affected less than 0.5% of RTX 5090, RTX 5090D, and RTX 5070 Ti and could reduce graphic performance by about 4%.

The problem has been fixed in new batches, but when buying an early or used RTX 5090, it is still worth checking GPU-Z: the program should show 176 ROPs.

Conclusion

The RTX 5090 proves itself in tasks where the RTX 4090 is already lacking: demanding 4K with path tracing, local AI, and 3D scenes that can utilize more than 24 GB of VRAM.

In regular games, the difference is much more modest. With an increase in TGP from 450 to 575 W and a starting price rise from $1599 to $1999, the advantage over the RTX 4090 without frame generation typically amounts to around 25-30%.

The advantages of the RTX 5090 are not limited to FPS: it features 32 GB of GDDR7, nearly 1.8 TB/s memory bandwidth, new RT and Tensor blocks, FP4, three NVENCs, and Multi Frame Generation.

For an RTX 4090 owner, replacing the card solely for improved standard gaming FPS is difficult to justify. However, in a new system for demanding 4K, local AI, or workloads that need more than 24 GB of VRAM, the advantages of the RTX 5090 become much clearer.

Overpaying for the expensive version of the RTX 5090 solely for factory overclocking is almost pointless: the difference with the Founders Edition typically amounts to only 3-5%. Specific models should be chosen based on cooling, noise levels, size, and price.

Basic

Label Name
NVIDIA
Platform
Desktop
Launch Date
January 2025
GPU Lithography
TSMC 4N
Model Name
GeForce RTX 5090
Generation
GeForce 50
Base Clock
2010 MHz
Boost Clock
2407 MHz
Bus Interface
PCIe Gen 5
Transistors
92.2
RT Cores
170, 4th Gen
Tensor Cores
?
Tensor Cores are specialized processing units designed specifically for deep learning, providing higher training and inference performance compared to FP32 training. They enable rapid computations in areas such as computer vision, natural language processing, speech recognition, text-to-speech conversion, and personalized recommendations. The two most notable applications of Tensor Cores are DLSS (Deep Learning Super Sampling) and AI Denoiser for noise reduction.
680, 5th Gen
TMUs
?
Texture Mapping Units (TMUs) serve as components of the GPU, which are capable of rotating, scaling, and distorting binary images, and then placing them as textures onto any plane of a given 3D model. This process is called texture mapping.
680
Foundry
TSMC
Architecture
NVIDIA Blackwell / GB202

Memory Specifications

Memory Size
32 GB GDDR7
Memory Type
GDDR7
Memory Bus
?
The memory bus width refers to the number of bits of data that the video memory can transfer within a single clock cycle. The larger the bus width, the greater the amount of data that can be transmitted instantaneously, making it one of the crucial parameters of video memory. The memory bandwidth is calculated as: Memory Bandwidth = Memory Frequency x Memory Bus Width / 8. Therefore, when the memory frequencies are similar, the memory bus width will determine the size of the memory bandwidth.
512-bit
Bandwidth
?
Memory bandwidth refers to the data transfer rate between the graphics chip and the video memory. It is measured in bytes per second, and the formula to calculate it is: memory bandwidth = working frequency × memory bus width / 8 bits.
1792 GB/s

Display and Media

Outputs
1x HDMI 2.1b
3x DisplayPort 2.1b with UHBR20

Theoretical Performance

Pixel Rate
?
Pixel fill rate refers to the number of pixels a graphics processing unit (GPU) can render per second, measured in MPixels/s (million pixels per second) or GPixels/s (billion pixels per second). It is the most commonly used metric to evaluate the pixel processing performance of a graphics card.
423.6 GPixel/s
Texture Rate
?
Texture fill rate refers to the number of texture map elements (texels) that a GPU can map to pixels in a single second.
1636.8 GTexel/s
FP16 (half)
?
An important metric for measuring GPU performance is floating-point computing capability. Half-precision floating-point numbers (16-bit) are used for applications like machine learning, where lower precision is acceptable. Single-precision floating-point numbers (32-bit) are used for common multimedia and graphics processing tasks, while double-precision floating-point numbers (64-bit) are required for scientific computing that demands a wide numeric range and high accuracy.
104.8 TFLOPS
FP64 (double)
?
An important metric for measuring GPU performance is floating-point computing capability. Double-precision floating-point numbers (64-bit) are required for scientific computing that demands a wide numeric range and high accuracy, while single-precision floating-point numbers (32-bit) are used for common multimedia and graphics processing tasks. Half-precision floating-point numbers (16-bit) are used for applications like machine learning, where lower precision is acceptable.
1.613 TFLOPS
FP32 (float)
?
An important metric for measuring GPU performance is floating-point computing capability. Single-precision floating-point numbers (32-bit) are used for common multimedia and graphics processing tasks, while double-precision floating-point numbers (64-bit) are required for scientific computing that demands a wide numeric range and high accuracy. Half-precision floating-point numbers (16-bit) are used for applications like machine learning, where lower precision is acceptable.
104.8 TFLOPS

Miscellaneous

SM Count
?
Multiple Streaming Processors (SPs), along with other resources, form a Streaming Multiprocessor (SM), which is also referred to as a GPU's major core. These additional resources include components such as warp schedulers, registers, and shared memory. The SM can be considered the heart of the GPU, similar to a CPU core, with registers and shared memory being scarce resources within the SM.
170
Shading Units
?
The most fundamental processing unit is the Streaming Processor (SP), where specific instructions and tasks are executed. GPUs perform parallel computing, which means multiple SPs work simultaneously to process tasks.
21760
L1 Cache
128 KB (per SM)
L2 Cache
96 MB
TDP
575 W TGP
Vulkan Version
?
Vulkan is a cross-platform graphics and compute API by Khronos Group, offering high performance and low CPU overhead. It lets developers control the GPU directly, reduces rendering overhead, and supports multi-threading and multi-core processors.
1.4
OpenCL Version
3.0
OpenGL
4.6
CUDA
CUDA Compute Capability 12.0
DirectX
12 Ultimate (12_2)
Power Connectors
1x 16-pin (12V-2x6)
ROPs
?
The Raster Operations Pipeline (ROPs) is primarily responsible for handling lighting and reflection calculations in games, as well as managing effects like anti-aliasing (AA), high resolution, smoke, and fire. The more demanding the anti-aliasing and lighting effects in a game, the higher the performance requirements for the ROPs; otherwise, it may result in a sharp drop in frame rate.
176
Shader Model
6.7
Suggested PSU
1000 W

Benchmarks

FP32 (float)
Score
104.8 TFLOPS
3DMark Steel Nomad
Score
14544
Blender
Score
15026.3
Vulkan
Score
366095
OpenCL
Score
368974

Compared to Other GPU

FP32 (float) / TFLOPS
163.4 +55.9%
105.525 +0.7%
79.478 -24.2%
65.572 -37.4%
3DMark Steel Nomad
14475 -0.5%
9237 -36.5%
8858 -39.1%
Blender
3548 -76.4%
1265.43 -91.6%
630 -95.8%
Vulkan
382809 +4.6%
100987 -72.4%
76392 -79.1%
49804 -86.4%
OpenCL
388405 +5.3%
126692 -65.7%
90580 -75.5%
66428 -82%