Die richtige GPU ist entscheidend für lokales AI Development. Dieser Guide vergleicht Consumer, Professional und Cloud GPUs.
Schnell-Überblick
| GPU | VRAM | Speed | Power | Price | Best For |
|---|---|---|---|---|---|
| RTX 4090 | 24GB | Schnell | 450W | $1,600 | Development, Inference |
| RTX 5090 | 32GB | Sehr Schnell | 575W | $1,999 | Development, Benchmarks |
| A100 | 80GB | Sehr Schnell | 250W | $20k | Training, Enterprise |
| H100 | 80GB | Blitzschnell | 700W | $30k+ | Large Training, Research |
| Cloud (RunPod) | Variabel | On-Demand | N/A | $0.76+/h | No Hardware Invest |
Consumer GPUs
RTX 4090 (Ada Lovelace)
Specs:
- VRAM: 24GB GDDR6X
- CUDA Cores: 16,384
- Tensor Cores: 512 (4th Gen)
- Memory Bandwidth: 960 GB/s
- TDP: 450W
- MSRP: $1,600 (Street: $1,400-2,000)
Models that Fit (Inference):
- Llama 3.1 70B (8-bit, needs 140GB): ❌ Won't Fit
- Llama 4 70B (4-bit quantized): ✅ Fits (20-24GB)
- Mistral 7B: ✅ Fits (14GB)
- Qwen 2.5 14B: ✅ Fits (18GB)
- GPT-3 Scale (175B): ❌ Needs 2× RTX 4090
Benchmarks:
- Model Inference Speed: 100-200 tokens/sec (4-bit models)
- GPU Utilization: ~80-95% during inference
- Noise: ~70dB (can hear it)
Power & Cooling:
- Braucht good Power Supply (850W minimum)
- Braucht good Airflow (avoid Mini-ITX cases)
- Can run multiple tasks (2 parallel inference jobs)
Best Use:
- Development & Prototyping
- Personal AI Projects
- Model Fine-Tuning (under 70B)
- Inference Server (1-2 concurrent users)
RTX 5090 (Blackwell, März 2026)
Specs:
- VRAM: 32GB GDDR7 (42% mehr als 4090)
- CUDA Cores: 21,760 (33% mehr)
- Memory Bandwidth: 1,792 GB/s
- TDP: 575W
- MSRP: $1,999
Performance vs RTX 4090:
- 27-35% faster
- 33% more VRAM
- 27% more expensive
- Probably not worth it for most (diminishing returns)
When Worth It:
- Running 72B models in 8-bit (needs 144GB)
- If 4090 can't handle your specific model
Best Use: Basically RTX 4090, but with more headroom
Professional GPUs
A100 (Professional, older but cheap used)
Specs:
- VRAM: 40GB or 80GB
- Tensor Performance: 312 TFLOPS (FP32)
- Memory Bandwidth: 2TB/s
- Power: 250W
- Price: $5-10k (used), $20k+ (new)
vs RTX 4090:
- 3× VRAM (40-80GB)
- Mais für Training optimiert (better precision)
- Viel teurer
- Weniger für inference performance
Best Use:
- Model Training
- Research
- Company/Lab Infrastructure
H100 (Latest Datacentre)
Specs:
- VRAM: 80GB HBM3
- Tensor Performance: 1,000+ TFLOPS (mit Sparsity)
- Memory Bandwidth: 3TB/s
- Power: 700W
- Price: $30k+ (new), $15-20k (used)
Models that Fit:
- Llama 4 405B (8-bit): ✅ Fits! (150GB)
- DeepSeek V3: ✅ Fits! (96GB)
- Training Large Models: ✅ Best Option
Best Use:
- Large-Scale Training
- Research Labs
- Cutting-Edge Models
- Enterprise ML Pipelines
Cloud GPU Pricing 2026
Selbst-Gehostete GPU Clouds (am günstigsten)
| Provider | RTX 4090 | A100 | H100 |
|---|---|---|---|
| RunPod Spot | $0.18-0.30/h | $1.50/h | $3.00/h |
| VAST.ai | $0.16-0.35/h | $0.80-2/h | $2-4/h |
| Lambda Labs | $0.36/h | $2.30/h | $4.50/h |
| Google Cloud | N/A | $3.67/h | $8+/h |
| AWS EC2 | N/A | $4.08/h | $9.08/h |
Best Value: RunPod Spot (bis zu 50% Rabatt, dafür preemptible)
Cost per Model Inference
Beispiel: Llama 4 70B Inference
RTX 4090 (owned):
- Hardware: $1,600 (amortized over 3 years = $44/Mo)
- Power: $20/Mo
- Hosting: $0 (home)
Total: ~$64/Mo or $0.003/inference
Cloud (RunPod Spot):
- Spot: $0.20/h
- 100 inferences/hour = $0.002/inference
Total: ~$0.002/inference
WINNER: Cloud is cheaper for low-volume!
Break-Even Analysis
Bei welcher Nutzung lohnt sich Hardware?
RTX 4090 Cost: $1,600 + $240/Year (Power) = $1,840
Cloud Cost per Hour: $0.25/h (RTX 4090 on RunPod)
Break-even:
$1,840 ÷ $0.25/h = 7,360 hours
= 307 days continuous use
= ~25 days at 24/7 usage
= ~3-5 months of heavy development
CONCLUSION: Buy hardware if >5 months active development
VRAM Requirements by Model & Precision
| Model | FP32 (Full) | FP16 (Half) | INT8 (8-bit) | INT4 (4-bit) |
|---|---|---|---|---|
| Llama 4 70B | 280GB | 140GB | 70GB | 20GB |
| Mistral 123B | 500GB | 250GB | 125GB | 40GB |
| Qwen 2.5 72B | 288GB | 144GB | 72GB | 22GB |
| Phi-4 14B | 56GB | 28GB | 14GB | 6GB |
| Gemma 2 9B | 36GB | 18GB | 9GB | 5GB |
Praktisch:
- RTX 4090 (24GB): 4-bit models bis ~70B
- A100 (80GB): 8-bit models bis ~400B
- H100 (80GB): Theoretisch ~600B mit MOE
Speed Comparisons
Token Generation Speed (Higher = Better)
Model: Llama 4 70B (4-bit quantized)
RTX 4090: 120-180 tokens/sec
A100: 150-220 tokens/sec
H100: 300-500 tokens/sec
Cloud (RunPod A100): 100-150 tokens/sec (network latency)
For Inference: Unterschied ist klein (all > 100 tok/sec = ausreichend)
Power & Cooling
| GPU | Power Draw | Power Supply | Cooling | Noise |
|---|---|---|---|---|
| RTX 4090 | 450W | 850W PSU | Good Case Fan | 70dB |
| RTX 5090 | 575W | 1000W PSU | Very Good | 75dB |
| A100 | 250W | Standard | Liquid (Data Ctr) | Data Ctr |
| H100 | 700W | 1200W PSU | Liquid Only | Data Ctr |
Home Setup mit RTX 4090: Kühlung ist wichtig! 70dB ist Lärm.
Praktische Szenarien
Szenario #1: Hobbyist mit Budget
Anforderungen:
- Llama 7B Chatbot lokal
- Budget: $500-1000
- Zuhause Setup
Beste Wahl: Gebrauchte RTX 3060 ($200) oder RTX 4070 ($600)
ollama pull llama2:7b
ollama run llama2:7b
Total Cost: $200-600 einmalig
Szenario #2: Developer mit Home Setup
Anforderungen:
- Llama 70B für Development
- Budget: $1500-2000
- Home Workstation
Beste Wahl: RTX 4090 ($1600)
# Runs locally on RTX 4090
ollama pull llama4:70b # 4-bit quantized
ollama run llama4:70b
Total Cost: $1,600 einmalig + $20/Mo Power
Szenario #3: Startup ohne Cloud Invest
Anforderungen:
- Llama 70B + GPT-4 Integration
- Multiple Users (5-10)
- Datenschutz (on-prem)
Beste Wahl: 2× RTX 4090 + Load Balancer
RTX 4090 #1 (User A, B, C)
RTX 4090 #2 (User D, E, F)
Nginx Load Balancer
Total: $3,200 + Setup
Alternative: Cloud A100 ($2.30/h) ist oft billiger!
Szenario #4: Large-Scale Production
Anforderungen:
- 1000+ concurrent Users
- Very Low Latency
- Enterprise SLA
Beste Wahl: Kubernetes on H100s
Kubernetes Cluster:
- 4× H100 GPUs
- Distributed Inference (vLLM)
- Auto-scaling
- Multi-Model Support
Cost: $500-2000/Month AWS + $30k for 4 H100s
Empfohlene Hardware Setups
| Use Case | GPU | RAM | Disk | Total Cost |
|---|---|---|---|---|
| Learning | RTX 3060 | 16GB | 500GB | $200-300 |
| Development | RTX 4090 | 32GB | 1TB SSD | $2,000-2,500 |
| Research | 2× RTX 4090 | 64GB | 2TB SSD | $4,000-4,500 |
| Production | Cloud H100 | N/A | N/A | $500-2000/Mo |
Top-5 Häufige Fehler
Problem #1: "RTX 4090 ist zu teuer"
- Lösung: Nutze Cloud GPU ($0.20/h) statt $1600 kauf
- Besser wenn: <100 Stunden/Monat Nutzung
Problem #2: "Mein RTX 4090 is bottlenecked by CPU"
- Ursache: Schwacher CPU oder alte Plattform
- Lösung: CPU sollte mind. 8 Cores, DDR5 RAM sein
Problem #3: "Cooling ist problematisch"
- Ursache: Bad Case oder verschmutzter Kühler
- Lösung: Bessere Airflow, regelmäßig saubern
Problem #4: "Cloud GPU ist teuer"
- Ursache: Langzeit-Jobs laufen 24/7
- Lösung: Optimiere inference, nutze batch processing
Problem #5: "Ich weiß nicht welche GPU ich brauche"
- Lösung: Start mit Cloud trial ($5 credits auf RunPod), dann entscheide
Zusätzliche Tipps
- Used Hardware: eBay/Amazon RTX 4090 gebraucht oft $1000-1200
- Cooling Mods: Aftermarket Cooler (ARCTIC, Alphacool) hilft
- Cloud Trial: Alle Clouds haben Free Credits ($5-100)
- Multi-GPU: 2× RTX 4090 Tensor Parallel ist fast wie 1× H100
Ressourcen
Letzte Aktualisierung: 21.03.2026
