Can you run mradermacher/blossom-v2-llama2-7b-i1-GGUF on RTX 4090 24GB?

mradermacher/blossom-v2-llama2-7b-i1-GGUF needs ~5.95 GB, which is within the 24 GB on RTX 4090 24GB by memory capacity (~18.05 GB headroom). Computed from stated size and quant, not a runtime guarantee.

mradermacher/blossom-v2-llama2-7b-i1-GGUF needs ~5.95 GB, which is within the 24 GB on RTX 4090 24GB by memory capacity (~18.05 GB headroom). Computed from stated size and quant, not a runtime guarantee.

Memory-capacity estimate across common GPUs

GPUVerdictVRAMNeeds
8GB laptop (no dedicated GPU)Fits (by memory)8 GB~5.95 GB
16GB laptop (no dedicated GPU)Fits (by memory)16 GB~5.95 GB
NVIDIA T4 16GB (free Colab)Fits (by memory)16 GB~5.95 GB
NVIDIA L4 24GBFits (by memory)24 GB~5.95 GB
RTX 3060 12GBFits (by memory)12 GB~5.95 GB
RTX 4080 16GBFits (by memory)16 GB~5.95 GB
RTX 3090 24GBFits (by memory)24 GB~5.95 GB
RTX 4090 24GBFits (by memory)24 GB~5.95 GB
A100 40GBFits (by memory)40 GB~5.95 GB
A100 80GBFits (by memory)80 GB~5.95 GB
H100 80GBFits (by memory)80 GB~5.95 GB
Apple M-series 16GB (unified)Fits (by memory)16 GB~5.95 GB
Apple M-series 32GB (unified)Fits (by memory)32 GB~5.95 GB
Apple M-series 64GB (unified)Fits (by memory)64 GB~5.95 GB
Apple M-series 128GB (unified)Fits (by memory)128 GB~5.95 GB
CPU / 32GB system RAMFits (by memory)32 GB~5.95 GB
CPU / 64GB system RAMFits (by memory)64 GB~5.95 GB

How this was computed

  • Weights + KV cache at the stated context + a runtime/overhead allowance; memory capacity only, not a benchmark.

Next steps

Methodology: weights + KV-cache at 8192 tokens + a runtime overhead allowance, from the recommended download size, parameter count, and quant. Computed from stated size and quant, not a runtime guarantee. Not a benchmark; runtimes differ.