The short answer first. The cheapest new GPU worth buying in 2026 is AMD's RX 9060 XT at its $299 MSRP; the best used deal for running local AI is a 24GB RTX 3090 at roughly $1,050; and the card most people should actually buy for both games and local models is the RTX 5060 Ti 16GB, which has been hovering around $560 new.
Buying a GPU this year is two decisions welded together: what you can afford, and what you want to run. The second one decides the first. For local AI, VRAM — not raw speed — is the spec that matters, and it is the reason a four-year-old used card can beat a brand-new mid-ranger.
What do GPUs cost in 2026?
Street prices in August 2026 are running well above the launch MSRPs, thanks to a DRAM and GDDR shortage that pushed memory prices up across the stack. The table below uses the lowest tracked new prices at major US retailers (Newegg, Amazon, B&H) in the first week of August 2026 — think of these as the realistic "you can actually buy one today" numbers.
| Card | VRAM | Launch MSRP | Street price |
|---|---|---|---|
| RTX 5090 | 32GB GDDR7 | $1,999 | ~$3,700–4,160 |
| RTX 5080 | 16GB GDDR7 | $999 | ~$1,250–1,450 |
| RTX 5070 Ti | 16GB GDDR7 | $749 | ~$919–970 |
| RTX 5070 | 12GB GDDR7 | $549 | ~$599–640 |
| RTX 5060 Ti 16GB | 16GB GDDR7 | $429 | ~$560–569 |
| RTX 5060 Ti 8GB | 8GB GDDR7 | $379 | ~$369–375 |
| RTX 5060 | 8GB GDDR7 | $299 | ~$339–370 |
| RX 9070 XT | 16GB GDDR6 | $599 | ~$549–689 |
| RX 9070 | 16GB GDDR6 | $549 | ~$554 |
| RX 9060 XT 16GB | 16GB GDDR6 | $349 | ~$332–349 |
| RX 9060 XT 8GB | 8GB GDDR6 | $299 | ~$299 |
| Intel Arc B580 | 12GB GDDR6 | $249 | ~$290 (deal) |
Two things stand out. The cheapest new card with a real shot at local AI is the RTX 5060 Ti 8GB at around $369, while the RX 9060 XT 16GB undercuts it with double the memory for about $349 — the best raw price-per-gigabyte on the new market. The RTX 5090, meanwhile, trades at nearly double its MSRP.
Why memory is the new spec war
The GDDR shortage did more than raise prices; it rewrote what reviewers mean by "value." Every card in the 8GB class lost its AI appeal overnight, while 16GB and 24GB cards gained a premium that has nothing to do with gaming performance. Industry reporting (Club386, Jan 2026) expects $100–200 of markup on mid-range cards to stick until memory production catches up.
The second force is bandwidth. Capacity decides what fits in memory; bandwidth decides how fast it runs. A 16GB card with a narrow bus will load a 32B model slower than a 24GB card with a wide one — which is why the used 24GB class keeps its price while shiny new 8GB cards depreciate.
Why is the used market the better value?
Used prices are tracked daily on sites like pcprice.watch, and they tell a different story. The card the local-AI community refuses to let die is the RTX 3090: 24GB of VRAM, plenty for 32B models at Q4 quantization, selling used for around $1,050 in July and August 2026 (pcprice.watch, Aug 2026). That is the same memory capacity as an $8,000+ workstation card, at a tenth of the price.
| Card | VRAM | Used price |
|---|---|---|
| RTX 3090 | 24GB | ~$1,050 |
| RTX 4090 | 24GB | ~$2,350 |
| RTX 5070 Ti | 16GB | ~$898 |
| RTX 5060 Ti | 8GB | ~$351 |
| RTX 5060 | 8GB | ~$319 |
| RTX 4080 | 16GB | ~€804 |
| RX 9070 XT | 16GB | ~€623 |
| RX 9070 | 16GB | ~€524 |
| RX 7900 XTX | 24GB | ~€728 |
| RX 7900 XT | 20GB | ~€744 |
The pattern is simple: 24GB-class cards hold their value because AI buyers want them; 8GB and 12GB cards are cheap because the market has outgrown them. If you are shopping used for AI, buy the memory, not the generation.
How much VRAM do you need to run AI?
The rule of thumb every local-AI guide converges on: a model needs roughly 1.2GB of VRAM per billion parameters at Q4 quantization — the standard "good enough" compression used by Ollama and llama.cpp (Hugging Face model docs, 2026). A 7B model fits in 8GB. A 14B model needs 12GB. A 24B model needs 16GB. A 32B model needs 24GB. A 70B model needs 48GB. Memory, not compute, is the ceiling.
| VRAM | What runs comfortably | Models it unlocks |
|---|---|---|
| 8GB | 7–8B chat + coding | Llama 3.1 8B, Qwen3-8B, DeepSeek R1 8B |
| 12GB | 12–14B | Qwen3-14B, Gemma 3 12B, Mistral 7B |
| 16GB | 24–27B | Mistral Small 3.2 24B, Qwen3-30B-A3B |
| 24GB | 32B, or 70B with heavy offload | Qwen3-32B, DeepSeek R1 32B |
| 48GB | 70B fully in VRAM | Llama 3.3 70B, DeepSeek R1 70B |
| 150GB+ | 200B-class MoE | Qwen3-235B-A22B (~150GB), DeepSeek R1 671B (~424GB) |
VRAM Calculator — What GPU Runs Your Model?
| GPU | VRAM | Fits | Headroom | Est. Speed | Price |
|---|---|---|---|---|---|
| RTX 4060 (8GB) | 8 GB | ❌ No | -16 GB | 3 tok/s | $280 |
| RTX 5060 Ti 8GB | 8 GB | ❌ No | -16 GB | 3 tok/s | $369 |
| RX 9060 XT 8GB | 8 GB | ❌ No | -16 GB | 3 tok/s | $299 |
| RTX 4070 Ti Super (16GB) | 16 GB | ❌ No | -8 GB | 6 tok/s | $550 |
| RTX 5060 Ti 16GB | 16 GB | ❌ No | -8 GB | 6 tok/s | $560 |
| RX 9060 XT 16GB | 16 GB | ❌ No | -8 GB | 6 tok/s | $349 |
| RTX 3090 24GB (used) | 24 GB | ✅ Yes | +0 GB | 8 tok/s | $1050 |
| RTX 4090 24GB (used) | 24 GB | ✅ Yes | +0 GB | 8 tok/s | $2350 |
| RTX 5090 32GB | 32 GB | ✅ Yes | +8 GB | 12 tok/s | $3700 |
| Mac Studio M4 Max 64GB | 64 GB | ✅ Yes | +40 GB | 9 tok/s | $3200 |
| Mac Studio M4 Max 128GB | 128 GB | ✅ Yes | +104 GB | 9 tok/s | $5000 |
| Mac Studio M4 Ultra 192GB | 192 GB | ✅ Yes | +168 GB | 9 tok/s | $8000 |
| RTX Pro 6000 96GB | 96 GB | ✅ Yes | +72 GB | 30 tok/s | $8000 |
| H100 80GB (cloud/hr) | 80 GB | ✅ Yes | +56 GB | 60 tok/s | $2/hr |
| H200 141GB (cloud/hr) | 141 GB | ✅ Yes | +117 GB | 60 tok/s | $3.5/hr |
Formula: params (B) × quantization factor × 1.2 safety + context overhead. Speed estimates for llama.cpp on consumer hardware. Cloud pricing per hour. Prices are August 2026 snapshots.
That map is why the buying advice sounds backwards: a used 24GB card beats a shiny new 8GB card for AI every single time. Long context changes the math too — a 32K-token window adds 2–4GB on top of the weights, so size one step up if you feed models long documents.
Can you run Claude on your own GPU?
Here is the clickbait, defused. You cannot run Claude on your own GPU, because Anthropic does not release Claude's weights. Claude is only available through the API, Amazon Bedrock, or Google Vertex AI — the "Claude Code" tool you can run locally is a thin client that still sends every request to the cloud.
If you want frontier-class quality at home, the open-weight stand-in is DeepSeek R1. At Q4 quantization it needs about 424GB of memory (DeepSeek model card, 2026), which is why it will never run on a desktop card. The practical ceiling for consumer hardware is the 32B class — models like Qwen3-32B that genuinely rival the cloud on everyday tasks, for the price of a used GPU.
The honest buying ladder looks like this. At 8GB you get a capable chat assistant. At 12–16GB you get a serious coding sidekick and agent scaffolding. At 24GB, a used 3090 gets you near-frontier quality for less than the price of a mid-range phone. Beyond that, you are in workstation or cloud territory.
Which GPU should you actually buy?
- Strict budget, games only: RX 9060 XT 8GB or Intel Arc B580 — under $300, 1080p capable, 8–12GB of memory.
- Best value, games + light AI: RX 9060 XT 16GB (~$349) or RTX 5060 Ti 8GB (~$369).
- Best AI value on a new card: RTX 5060 Ti 16GB (~$560) — the cheapest new card with room for 24B models.
- Best AI value, period: a used RTX 3090 24GB (~$1,050) for 32B models at near-frontier quality.
- No GPU needed: rent one in the cloud for a few dollars an hour and skip the upgrade cycle entirely.
Sources and further reading
Prices were verified against live US retailer listings and the pcprice.watch used-price tracker in the first week of August 2026. GPU pricing moves weekly in this shortage, so treat every number as a snapshot, not a promise.
The Best AI Models of 2026, Ranked by Real Users
Open Source Won the Model War. Here's What Comes Next
Small Language Models Punch Way Above Their Weight
AI Agents Hit the Payroll — What Works and What's a Demo
NVIDIA GeForce RTX 50-series specs
pcprice.watch — used GPU price tracker
Ollama library — the easiest way to run local models
Cover photo: RTX 3090 Founders Edition, Wikimedia Commons
About Savviest — editorial policy and methods
Buy for the model you actually run, not the marketing budget of the card. In 2026 that usually means a $369 RTX 5060 Ti, a used $1,050 RTX 3090, or — if you genuinely need frontier scale — a cloud instance. The GPU that fits your model is the GPU worth your money.
How much VRAM do I need for a 70B model?
At Q4 quantization (the standard "good enough" compression), a 70B model needs roughly 48GB of VRAM to run fully in memory. At Q8 it needs ~70GB. With heavy CPU offloading you can run 70B on 24GB, but expect 2-5 tokens/sec. For usable speed, 48GB+ (dual 24GB cards or workstation) is the floor.
Is a used RTX 3090 still the best value in 2026?
Yes, for pure AI workloads. At ~$1,050 for 24GB VRAM, it delivers the best dollars-per-gigabyte on the market. The trade-offs: no warranty, higher power draw (350W+), no NVLink on consumer cards (so no multi-GPU pooling), and aging architecture. For gaming + AI, the RTX 5060 Ti 16GB at ~$560 new is the better all-rounder.
Can I run local AI on a Mac instead of a GPU?
Yes. Apple Silicon uses unified memory — an M4 Max with 64GB runs Llama 3 70B at Q4-Q6; 128GB runs it at FP16. M4 Ultra (192GB) runs DeepSeek R1 671B. The trade-off: 2-3x slower tokens/sec than equivalent NVIDIA, but unmatched memory capacity per dollar. For privacy-first or memory-bound workloads, Mac wins.
What about AMD ROCm — is it viable for AI in 2026?
Improving but still behind CUDA. RX 7900 XTX (24GB) at ~$899 is the best AMD value. ROCm 6.x supports PyTorch, llama.cpp, and vLLM on Linux. Windows support is experimental. If you are Linux-comfortable and want 24GB under $1,000, it works. For plug-and-play, NVIDIA still wins.
Should I buy or rent GPU compute in 2026?
Buy if you run models daily (utilization >40%). A used RTX 3090 at $1,050 breaks even vs. cloud H100 at $2/hr in ~500 hours. Rent if workloads are bursty, experimental, or need >48GB VRAM. Cloud also wins for training — H100/H200 bandwidth and NVLink make multi-GPU training viable.
Sources and further reading
Prices were verified against live US retailer listings and the pcprice.watch used-price tracker in the first week of August 2026. GPU pricing moves weekly in this shortage, so treat every number as a snapshot, not a promise.
The Best AI Models of 2026, Ranked by Real Users
Open Source Won the Model War. Here's What Comes Next
Small Language Models Punch Way Above Their Weight
AI Agents Hit the Payroll — What Works and What's a Demo
NVIDIA GeForce RTX 50-series specs
pcprice.watch — used GPU price tracker
Ollama library — the easiest way to run local models
Cover photo: RTX 3090 Founders Edition, Wikimedia Commons
About Savviest — editorial policy and methods
Bottom line
Prices were verified against live US retailer listings and the pcprice.watch used-price tracker in the first week of August 2026. GPU pricing moves weekly in this shortage, so treat every number as a snapshot, not a promise.
What we still don't know
This is a fast-moving story. We update the post as new facts land — and we'll flag it when we do.
Enjoyed this? Pay it forward
Five people forward this newsletter before they finish their coffee. Make it six.



