Skip to main content

NVIDIA RTX 5090 (Blackwell Consumer Flagship)

Product Overview​

The NVIDIA RTX 5090, unveiled at CES 2025-01, is the consumer flagship bringing the Blackwell architecture to consumer GPUs for the first time. With 32GB GDDR7 memory, 21,760 CUDA cores, and a 575W TDP, it delivers 3,352 TOPS of AI compute (FP4) — 2.5× that of the RTX 4090.

Positioned for local LLM inference (70B+ models), Stable Diffusion XL training, and consumer AI developers.

Core Specifications​

ParameterValue
ArchitectureBlackwell (GB202)
Process NodeTSMC 4N (custom 5nm)
CUDA Cores21,760
Tensor Cores680 (5th Gen)
RT Cores170 (4th Gen)
Base Clock2.01 GHz
Boost Clock2.41 GHz
Memory32 GB GDDR7
Memory Bandwidth1,792 GB/s (28 Gbps × 512-bit)
FP32 Compute104.8 TFLOPS
FP16 Tensor419 TFLOPS (sparse)
FP8 Tensor838 TFLOPS (sparse)
FP4 Tensor3,352 TOPS (sparse)
INT8 Tensor1,676 TOPS
TDP575 W
Power Connector1× 16-pin (12V-2x6)
MSRP$1,999
Launch Date2025-01-30

RTX 5090 vs RTX 4090 Comparison​

MetricRTX 5090RTX 4090Improvement
ArchitectureBlackwellAda LovelaceNew gen
CUDA Cores21,76016,3841.33×
Memory32GB GDDR724GB GDDR6X1.33×
Memory Bandwidth1,792 GB/s1,008 GB/s1.78×
FP16 Tensor419 TFLOPS165 TFLOPS2.5×
FP4 Tensor3,352 TOPSN/ANew
TDP575W450W1.28×
Price$1,999$1,5991.25×

Blackwell New Features​

FP4 Precision Support​

  • Native FP4 Tensor Cores (first time on consumer GPUs).
  • Reduces inference memory footprint by 50% (vs FP8).
  • 70B LLM can run FP4 quantized within 32GB memory (~40GB model compressed).

DLSS 4 Multi Frame Generation​

  • Multi Frame Generation: Generates 3 frames from 1 (vs DLSS 3's 1 frame from 1).
  • Gaming-only, but showcases Blackwell's compute power.

GDDR7 Memory​

  • 28 Gbps speed (vs GDDR6X 21 Gbps).
  • 1,792 GB/s bandwidth = 2× RTX 4090.
  • Alleviates the memory-bound bottleneck in LLM inference.

LLM Inference Performance​

ModelQuantizationRTX 5090 (32GB)RTX 4090 (24GB)Improvement
Llama 3 8BFP16~95 tok/s~70 tok/s1.36×
Llama 3 70BFP4~28 tok/sOOMBreakthrough
Llama 3 70BINT4~22 tok/s~15 tok/s1.47×
Mixtral 8x7BINT4~45 tok/s~32 tok/s1.41×
Qwen 2.5 72BFP4~26 tok/sOOMBreakthrough

70B model FP4 quantized (~40GB) fully fits in VRAM — 32GB memory is the key enabler.

Vendor Information​

ParameterValue
VendorNVIDIA Corporation
Product Pagehttps://www.nvidia.com/en-us/geforce/graphics-cards/50-series/rtx-5090/
MSRP$1,999 (FE Founders Edition)
Target MarketConsumer AI, creators, researchers, local LLM

Use Cases​

  • ✅ Local 70B LLM inference (FP4 quantized, 32GB VRAM)
  • ✅ Stable Diffusion XL / Flux training and inference
  • ✅ Video production (DaVinci Resolve AI acceleration)
  • ✅ 8K gaming + frame generation
  • ❌ Data center (use H100/B200 instead)
  • ❌ Multi-node training (lacks NVLink)