Die richtige GPU ist entscheidend für lokales AI Development. Dieser Guide vergleicht Consumer, Professional und Cloud GPUs.

Schnell-Überblick

GPU VRAM Speed Power Price Best For
RTX 4090 24GB Schnell 450W $1,600 Development, Inference
RTX 5090 32GB Sehr Schnell 575W $1,999 Development, Benchmarks
A100 80GB Sehr Schnell 250W $20k Training, Enterprise
H100 80GB Blitzschnell 700W $30k+ Large Training, Research
Cloud (RunPod) Variabel On-Demand N/A $0.76+/h No Hardware Invest

Consumer GPUs

RTX 4090 (Ada Lovelace)

Specs:

  • VRAM: 24GB GDDR6X
  • CUDA Cores: 16,384
  • Tensor Cores: 512 (4th Gen)
  • Memory Bandwidth: 960 GB/s
  • TDP: 450W
  • MSRP: $1,600 (Street: $1,400-2,000)

Models that Fit (Inference):

  • Llama 3.1 70B (8-bit, needs 140GB): ❌ Won't Fit
  • Llama 4 70B (4-bit quantized): ✅ Fits (20-24GB)
  • Mistral 7B: ✅ Fits (14GB)
  • Qwen 2.5 14B: ✅ Fits (18GB)
  • GPT-3 Scale (175B): ❌ Needs 2× RTX 4090

Benchmarks:

  • Model Inference Speed: 100-200 tokens/sec (4-bit models)
  • GPU Utilization: ~80-95% during inference
  • Noise: ~70dB (can hear it)

Power & Cooling:

  • Braucht good Power Supply (850W minimum)
  • Braucht good Airflow (avoid Mini-ITX cases)
  • Can run multiple tasks (2 parallel inference jobs)

Best Use:

  • Development & Prototyping
  • Personal AI Projects
  • Model Fine-Tuning (under 70B)
  • Inference Server (1-2 concurrent users)

RTX 5090 (Blackwell, März 2026)

Specs:

  • VRAM: 32GB GDDR7 (42% mehr als 4090)
  • CUDA Cores: 21,760 (33% mehr)
  • Memory Bandwidth: 1,792 GB/s
  • TDP: 575W
  • MSRP: $1,999

Performance vs RTX 4090:

  • 27-35% faster
  • 33% more VRAM
  • 27% more expensive
  • Probably not worth it for most (diminishing returns)

When Worth It:

  • Running 72B models in 8-bit (needs 144GB)
  • If 4090 can't handle your specific model

Best Use: Basically RTX 4090, but with more headroom

Professional GPUs

A100 (Professional, older but cheap used)

Specs:

  • VRAM: 40GB or 80GB
  • Tensor Performance: 312 TFLOPS (FP32)
  • Memory Bandwidth: 2TB/s
  • Power: 250W
  • Price: $5-10k (used), $20k+ (new)

vs RTX 4090:

  • 3× VRAM (40-80GB)
  • Mais für Training optimiert (better precision)
  • Viel teurer
  • Weniger für inference performance

Best Use:

  • Model Training
  • Research
  • Company/Lab Infrastructure

H100 (Latest Datacentre)

Specs:

  • VRAM: 80GB HBM3
  • Tensor Performance: 1,000+ TFLOPS (mit Sparsity)
  • Memory Bandwidth: 3TB/s
  • Power: 700W
  • Price: $30k+ (new), $15-20k (used)

Models that Fit:

  • Llama 4 405B (8-bit): ✅ Fits! (150GB)
  • DeepSeek V3: ✅ Fits! (96GB)
  • Training Large Models: ✅ Best Option

Best Use:

  • Large-Scale Training
  • Research Labs
  • Cutting-Edge Models
  • Enterprise ML Pipelines

Cloud GPU Pricing 2026

Selbst-Gehostete GPU Clouds (am günstigsten)

Provider RTX 4090 A100 H100
RunPod Spot $0.18-0.30/h $1.50/h $3.00/h
VAST.ai $0.16-0.35/h $0.80-2/h $2-4/h
Lambda Labs $0.36/h $2.30/h $4.50/h
Google Cloud N/A $3.67/h $8+/h
AWS EC2 N/A $4.08/h $9.08/h

Best Value: RunPod Spot (bis zu 50% Rabatt, dafür preemptible)

Cost per Model Inference

Beispiel: Llama 4 70B Inference

RTX 4090 (owned):
- Hardware: $1,600 (amortized over 3 years = $44/Mo)
- Power: $20/Mo
- Hosting: $0 (home)
Total: ~$64/Mo or $0.003/inference

Cloud (RunPod Spot):
- Spot: $0.20/h
- 100 inferences/hour = $0.002/inference
Total: ~$0.002/inference

WINNER: Cloud is cheaper for low-volume!

Break-Even Analysis

Bei welcher Nutzung lohnt sich Hardware?

RTX 4090 Cost: $1,600 + $240/Year (Power) = $1,840

Cloud Cost per Hour: $0.25/h (RTX 4090 on RunPod)

Break-even:
$1,840 ÷ $0.25/h = 7,360 hours
= 307 days continuous use
= ~25 days at 24/7 usage
= ~3-5 months of heavy development

CONCLUSION: Buy hardware if >5 months active development

VRAM Requirements by Model & Precision

Model FP32 (Full) FP16 (Half) INT8 (8-bit) INT4 (4-bit)
Llama 4 70B 280GB 140GB 70GB 20GB
Mistral 123B 500GB 250GB 125GB 40GB
Qwen 2.5 72B 288GB 144GB 72GB 22GB
Phi-4 14B 56GB 28GB 14GB 6GB
Gemma 2 9B 36GB 18GB 9GB 5GB

Praktisch:

  • RTX 4090 (24GB): 4-bit models bis ~70B
  • A100 (80GB): 8-bit models bis ~400B
  • H100 (80GB): Theoretisch ~600B mit MOE

Speed Comparisons

Token Generation Speed (Higher = Better)

Model: Llama 4 70B (4-bit quantized)

RTX 4090: 120-180 tokens/sec
A100: 150-220 tokens/sec
H100: 300-500 tokens/sec
Cloud (RunPod A100): 100-150 tokens/sec (network latency)

For Inference: Unterschied ist klein (all > 100 tok/sec = ausreichend)

Power & Cooling

GPU Power Draw Power Supply Cooling Noise
RTX 4090 450W 850W PSU Good Case Fan 70dB
RTX 5090 575W 1000W PSU Very Good 75dB
A100 250W Standard Liquid (Data Ctr) Data Ctr
H100 700W 1200W PSU Liquid Only Data Ctr

Home Setup mit RTX 4090: Kühlung ist wichtig! 70dB ist Lärm.

Praktische Szenarien

Szenario #1: Hobbyist mit Budget

Anforderungen:

  • Llama 7B Chatbot lokal
  • Budget: $500-1000
  • Zuhause Setup

Beste Wahl: Gebrauchte RTX 3060 ($200) oder RTX 4070 ($600)

ollama pull llama2:7b
ollama run llama2:7b

Total Cost: $200-600 einmalig

Szenario #2: Developer mit Home Setup

Anforderungen:

  • Llama 70B für Development
  • Budget: $1500-2000
  • Home Workstation

Beste Wahl: RTX 4090 ($1600)

# Runs locally on RTX 4090
ollama pull llama4:70b  # 4-bit quantized
ollama run llama4:70b

Total Cost: $1,600 einmalig + $20/Mo Power

Szenario #3: Startup ohne Cloud Invest

Anforderungen:

  • Llama 70B + GPT-4 Integration
  • Multiple Users (5-10)
  • Datenschutz (on-prem)

Beste Wahl: 2× RTX 4090 + Load Balancer

RTX 4090 #1 (User A, B, C)
RTX 4090 #2 (User D, E, F)
Nginx Load Balancer
Total: $3,200 + Setup

Alternative: Cloud A100 ($2.30/h) ist oft billiger!

Szenario #4: Large-Scale Production

Anforderungen:

  • 1000+ concurrent Users
  • Very Low Latency
  • Enterprise SLA

Beste Wahl: Kubernetes on H100s

Kubernetes Cluster:
- 4× H100 GPUs
- Distributed Inference (vLLM)
- Auto-scaling
- Multi-Model Support

Cost: $500-2000/Month AWS + $30k for 4 H100s

Empfohlene Hardware Setups

Use Case GPU RAM Disk Total Cost
Learning RTX 3060 16GB 500GB $200-300
Development RTX 4090 32GB 1TB SSD $2,000-2,500
Research 2× RTX 4090 64GB 2TB SSD $4,000-4,500
Production Cloud H100 N/A N/A $500-2000/Mo

Top-5 Häufige Fehler

Problem #1: "RTX 4090 ist zu teuer"

  • Lösung: Nutze Cloud GPU ($0.20/h) statt $1600 kauf
  • Besser wenn: <100 Stunden/Monat Nutzung

Problem #2: "Mein RTX 4090 is bottlenecked by CPU"

  • Ursache: Schwacher CPU oder alte Plattform
  • Lösung: CPU sollte mind. 8 Cores, DDR5 RAM sein

Problem #3: "Cooling ist problematisch"

  • Ursache: Bad Case oder verschmutzter Kühler
  • Lösung: Bessere Airflow, regelmäßig saubern

Problem #4: "Cloud GPU ist teuer"

  • Ursache: Langzeit-Jobs laufen 24/7
  • Lösung: Optimiere inference, nutze batch processing

Problem #5: "Ich weiß nicht welche GPU ich brauche"

  • Lösung: Start mit Cloud trial ($5 credits auf RunPod), dann entscheide

Zusätzliche Tipps

  • Used Hardware: eBay/Amazon RTX 4090 gebraucht oft $1000-1200
  • Cooling Mods: Aftermarket Cooler (ARCTIC, Alphacool) hilft
  • Cloud Trial: Alle Clouds haben Free Credits ($5-100)
  • Multi-GPU: 2× RTX 4090 Tensor Parallel ist fast wie 1× H100

Ressourcen

Letzte Aktualisierung: 21.03.2026