The short answer first. The cheapest new GPU worth buying in 2026 is AMD's RX 9060 XT at its $299 MSRP; the best used deal for running local AI is a 24GB RTX 3090 at roughly $1,050; and the card most people should actually buy for both games and local models is the RTX 5060 Ti 16GB, which has been hovering around $560 new.

Buying a GPU this year is two decisions welded together: what you can afford, and what you want to run. The second one decides the first. For local AI, VRAM — not raw speed — is the spec that matters, and it is the reason a four-year-old used card can beat a brand-new mid-ranger.

What do GPUs cost in 2026?

Street prices in August 2026 are running well above the launch MSRPs, thanks to a DRAM and GDDR shortage that pushed memory prices up across the stack. The table below uses the lowest tracked new prices at major US retailers (Newegg, Amazon, B&H) in the first week of August 2026 — think of these as the realistic "you can actually buy one today" numbers.

New GPU prices, US street, early Aug 2026
CardVRAMLaunch MSRPStreet price
RTX 509032GB GDDR7$1,999~$3,700–4,160
RTX 508016GB GDDR7$999~$1,250–1,450
RTX 5070 Ti16GB GDDR7$749~$919–970
RTX 507012GB GDDR7$549~$599–640
RTX 5060 Ti 16GB16GB GDDR7$429~$560–569
RTX 5060 Ti 8GB8GB GDDR7$379~$369–375
RTX 50608GB GDDR7$299~$339–370
RX 9070 XT16GB GDDR6$599~$549–689
RX 907016GB GDDR6$549~$554
RX 9060 XT 16GB16GB GDDR6$349~$332–349
RX 9060 XT 8GB8GB GDDR6$299~$299
Intel Arc B58012GB GDDR6$249~$290 (deal)

Two things stand out. The cheapest new card with a real shot at local AI is the RTX 5060 Ti 8GB at around $369, while the RX 9060 XT 16GB undercuts it with double the memory for about $349 — the best raw price-per-gigabyte on the new market. The RTX 5090, meanwhile, trades at nearly double its MSRP.

Why memory is the new spec war

The GDDR shortage did more than raise prices; it rewrote what reviewers mean by "value." Every card in the 8GB class lost its AI appeal overnight, while 16GB and 24GB cards gained a premium that has nothing to do with gaming performance. Industry reporting (Club386, Jan 2026) expects $100–200 of markup on mid-range cards to stick until memory production catches up.

The second force is bandwidth. Capacity decides what fits in memory; bandwidth decides how fast it runs. A 16GB card with a narrow bus will load a 32B model slower than a 24GB card with a wide one — which is why the used 24GB class keeps its price while shiny new 8GB cards depreciate.

Why is the used market the better value?

Used prices are tracked daily on sites like pcprice.watch, and they tell a different story. The card the local-AI community refuses to let die is the RTX 3090: 24GB of VRAM, plenty for 32B models at Q4 quantization, selling used for around $1,050 in July and August 2026 (pcprice.watch, Aug 2026). That is the same memory capacity as an $8,000+ workstation card, at a tenth of the price.

Used GPU prices, tracked Jul–Aug 2026
CardVRAMUsed price
RTX 309024GB~$1,050
RTX 409024GB~$2,350
RTX 5070 Ti16GB~$898
RTX 5060 Ti8GB~$351
RTX 50608GB~$319
RTX 408016GB~€804
RX 9070 XT16GB~€623
RX 907016GB~€524
RX 7900 XTX24GB~€728
RX 7900 XT20GB~€744
24GBThe VRAM class local-AI buyers actually pay for — a used RTX 3090 runs it for about · $1,050

The pattern is simple: 24GB-class cards hold their value because AI buyers want them; 8GB and 12GB cards are cheap because the market has outgrown them. If you are shopping used for AI, buy the memory, not the generation.

How much VRAM do you need to run AI?

The rule of thumb every local-AI guide converges on: a model needs roughly 1.2GB of VRAM per billion parameters at Q4 quantization — the standard "good enough" compression used by Ollama and llama.cpp (Hugging Face model docs, 2026). A 7B model fits in 8GB. A 14B model needs 12GB. A 24B model needs 16GB. A 32B model needs 24GB. A 70B model needs 48GB. Memory, not compute, is the ceiling.

The VRAM-to-model map, at Q4 quantization
VRAMWhat runs comfortablyModels it unlocks
8GB7–8B chat + codingLlama 3.1 8B, Qwen3-8B, DeepSeek R1 8B
12GB12–14BQwen3-14B, Gemma 3 12B, Mistral 7B
16GB24–27BMistral Small 3.2 24B, Qwen3-30B-A3B
24GB32B, or 70B with heavy offloadQwen3-32B, DeepSeek R1 32B
48GB70B fully in VRAMLlama 3.3 70B, DeepSeek R1 70B
150GB+200B-class MoEQwen3-235B-A22B (~150GB), DeepSeek R1 671B (~424GB)

VRAM Calculator — What GPU Runs Your Model?

Estimated VRAM Required
24 GB
Base: 16.0 GB + Context: 4 GB + 20% safety = 24 GB
Best Match
RTX 3090 24GB (used) (24 GB)
Headroom: +0 GB | Est. 8 tok/s | $1050
GPUVRAMFitsHeadroomEst. SpeedPrice
RTX 4060 (8GB)8 GB❌ No-16 GB3 tok/s$280
RTX 5060 Ti 8GB8 GB❌ No-16 GB3 tok/s$369
RX 9060 XT 8GB8 GB❌ No-16 GB3 tok/s$299
RTX 4070 Ti Super (16GB)16 GB❌ No-8 GB6 tok/s$550
RTX 5060 Ti 16GB16 GB❌ No-8 GB6 tok/s$560
RX 9060 XT 16GB16 GB❌ No-8 GB6 tok/s$349
RTX 3090 24GB (used)24 GB✅ Yes+0 GB8 tok/s$1050
RTX 4090 24GB (used)24 GB✅ Yes+0 GB8 tok/s$2350
RTX 5090 32GB32 GB✅ Yes+8 GB12 tok/s$3700
Mac Studio M4 Max 64GB64 GB✅ Yes+40 GB9 tok/s$3200
Mac Studio M4 Max 128GB128 GB✅ Yes+104 GB9 tok/s$5000
Mac Studio M4 Ultra 192GB192 GB✅ Yes+168 GB9 tok/s$8000
RTX Pro 6000 96GB96 GB✅ Yes+72 GB30 tok/s$8000
H100 80GB (cloud/hr)80 GB✅ Yes+56 GB60 tok/s$2/hr
H200 141GB (cloud/hr)141 GB✅ Yes+117 GB60 tok/s$3.5/hr

Formula: params (B) × quantization factor × 1.2 safety + context overhead. Speed estimates for llama.cpp on consumer hardware. Cloud pricing per hour. Prices are August 2026 snapshots.

That map is why the buying advice sounds backwards: a used 24GB card beats a shiny new 8GB card for AI every single time. Long context changes the math too — a 32K-token window adds 2–4GB on top of the weights, so size one step up if you feed models long documents.

Can you run Claude on your own GPU?

Here is the clickbait, defused. You cannot run Claude on your own GPU, because Anthropic does not release Claude's weights. Claude is only available through the API, Amazon Bedrock, or Google Vertex AI — the "Claude Code" tool you can run locally is a thin client that still sends every request to the cloud.

If you want frontier-class quality at home, the open-weight stand-in is DeepSeek R1. At Q4 quantization it needs about 424GB of memory (DeepSeek model card, 2026), which is why it will never run on a desktop card. The practical ceiling for consumer hardware is the 32B class — models like Qwen3-32B that genuinely rival the cloud on everyday tasks, for the price of a used GPU.

~424GBVRAM needed to run DeepSeek R1 671B at Q4 — the real price of frontier-class local AI

The honest buying ladder looks like this. At 8GB you get a capable chat assistant. At 12–16GB you get a serious coding sidekick and agent scaffolding. At 24GB, a used 3090 gets you near-frontier quality for less than the price of a mid-range phone. Beyond that, you are in workstation or cloud territory.

Which GPU should you actually buy?

  • Strict budget, games only: RX 9060 XT 8GB or Intel Arc B580 — under $300, 1080p capable, 8–12GB of memory.
  • Best value, games + light AI: RX 9060 XT 16GB (~$349) or RTX 5060 Ti 8GB (~$369).
  • Best AI value on a new card: RTX 5060 Ti 16GB (~$560) — the cheapest new card with room for 24B models.
  • Best AI value, period: a used RTX 3090 24GB (~$1,050) for 32B models at near-frontier quality.
  • No GPU needed: rent one in the cloud for a few dollars an hour and skip the upgrade cycle entirely.

Sources and further reading

Prices were verified against live US retailer listings and the pcprice.watch used-price tracker in the first week of August 2026. GPU pricing moves weekly in this shortage, so treat every number as a snapshot, not a promise.

The Best AI Models of 2026, Ranked by Real Users

Open Source Won the Model War. Here's What Comes Next

Small Language Models Punch Way Above Their Weight

AI Agents Hit the Payroll — What Works and What's a Demo

NVIDIA GeForce RTX 50-series specs

pcprice.watch — used GPU price tracker

Ollama library — the easiest way to run local models

Cover photo: RTX 3090 Founders Edition, Wikimedia Commons

About Savviest — editorial policy and methods

Buy for the model you actually run, not the marketing budget of the card. In 2026 that usually means a $369 RTX 5060 Ti, a used $1,050 RTX 3090, or — if you genuinely need frontier scale — a cloud instance. The GPU that fits your model is the GPU worth your money.

How much VRAM do I need for a 70B model?

At Q4 quantization (the standard "good enough" compression), a 70B model needs roughly 48GB of VRAM to run fully in memory. At Q8 it needs ~70GB. With heavy CPU offloading you can run 70B on 24GB, but expect 2-5 tokens/sec. For usable speed, 48GB+ (dual 24GB cards or workstation) is the floor.

Is a used RTX 3090 still the best value in 2026?

Yes, for pure AI workloads. At ~$1,050 for 24GB VRAM, it delivers the best dollars-per-gigabyte on the market. The trade-offs: no warranty, higher power draw (350W+), no NVLink on consumer cards (so no multi-GPU pooling), and aging architecture. For gaming + AI, the RTX 5060 Ti 16GB at ~$560 new is the better all-rounder.

Can I run local AI on a Mac instead of a GPU?

Yes. Apple Silicon uses unified memory — an M4 Max with 64GB runs Llama 3 70B at Q4-Q6; 128GB runs it at FP16. M4 Ultra (192GB) runs DeepSeek R1 671B. The trade-off: 2-3x slower tokens/sec than equivalent NVIDIA, but unmatched memory capacity per dollar. For privacy-first or memory-bound workloads, Mac wins.

What about AMD ROCm — is it viable for AI in 2026?

Improving but still behind CUDA. RX 7900 XTX (24GB) at ~$899 is the best AMD value. ROCm 6.x supports PyTorch, llama.cpp, and vLLM on Linux. Windows support is experimental. If you are Linux-comfortable and want 24GB under $1,000, it works. For plug-and-play, NVIDIA still wins.

Should I buy or rent GPU compute in 2026?

Buy if you run models daily (utilization >40%). A used RTX 3090 at $1,050 breaks even vs. cloud H100 at $2/hr in ~500 hours. Rent if workloads are bursty, experimental, or need >48GB VRAM. Cloud also wins for training — H100/H200 bandwidth and NVLink make multi-GPU training viable.

Sources and further reading

Prices were verified against live US retailer listings and the pcprice.watch used-price tracker in the first week of August 2026. GPU pricing moves weekly in this shortage, so treat every number as a snapshot, not a promise.

The Best AI Models of 2026, Ranked by Real Users

Open Source Won the Model War. Here's What Comes Next

Small Language Models Punch Way Above Their Weight

AI Agents Hit the Payroll — What Works and What's a Demo

NVIDIA GeForce RTX 50-series specs

pcprice.watch — used GPU price tracker

Ollama library — the easiest way to run local models

Cover photo: RTX 3090 Founders Edition, Wikimedia Commons

About Savviest — editorial policy and methods

Bottom line

Prices were verified against live US retailer listings and the pcprice.watch used-price tracker in the first week of August 2026. GPU pricing moves weekly in this shortage, so treat every number as a snapshot, not a promise.

What we still don't know

This is a fast-moving story. We update the post as new facts land — and we'll flag it when we do.

Enjoyed this? Pay it forward

Five people forward this newsletter before they finish their coffee. Make it six.

Read moreShare on X