LLM VRAM Calculator Widget

Before renting a GPU or buying one for local inference, developers can estimate whether a model fits. Enter the parameter count and precision, optionally the layers, hidden size and context for the key-value cache, and an overhead allowance, and get an estimate in GB with the weights, cache and overhead shown separately.

Developer Tools Calculator Runs in your browser Free · no ads

Customize your widget

Theme
Auto follows the visitor's light/dark setting.
Style
Attribution on your page
Optional and entirely your choice. The exact line is shown in the code below; it links to the tool with rel="nofollow".
More options
Starting values
Leave blank to use the widget's defaults. Visitors can still change every value.

Live preview

Exactly what your visitors will see

Embed code

<iframe src="https://a2z.tools/embed/w/llm-vram-calculator" title="LLM VRAM Calculator by A2Z Tools" width="100%" height="820" style="border:0;width:100%" loading="lazy" allow="clipboard-write"></iframe>

A plain iframe. Works everywhere, including site builders that strip scripts. Adjust height if your content needs more room.

Works with

How it works

The Hugging Face Transformers guide to optimising LLMs puts the arithmetic simply: loading a model with X billion parameters needs roughly 4X GB in float32 and 2X GB in bfloat16 or float16, so a 70B model needs about 140 GB at 16-bit. 8-bit quantisation stores one byte per parameter and 4-bit about half a byte. During generation the key-value cache adds memory that grows with context: the guide counts 2 x sequence length x layers x hidden size values, which for its octocoder example (40 layers, hidden size 6,144, 16,000 tokens) is 7,864,320,000 values - about 15.7 GB in float16. The widget multiplies that by the batch size and by the ratio of key-value heads to attention heads, which is below 1 for grouped-query attention models. An overhead percentage, your own allowance for the CUDA context, activations and fragmentation, is added on top. Results are in decimal GB with GiB alongside.

Calculation method

  • Weights (bytes) = parameters x bytes per parameter (FP32 4, FP16/BF16 2, INT8/FP8 1, 4-bit 0.5)
  • KV cache (bytes) = 2 x context length x layers x hidden size x batch x (KV heads / attention heads) x bytes per value
  • Total = (weights + KV cache) x (1 + overhead % / 100)
  • GB = bytes / 10^9; GiB = bytes / 2^30

Worked examples

7B model in FP16 with a 4k context

Inputs: 7B; FP16; 32 layers; hidden 4,096; 4,096 tokens; batch 1; KV ratio 1; 10% overhead

Result: 17.8 GB (16.5 GiB): weights 14 GB, KV cache 2.15 GB, overhead 1.61 GB

2 x 4,096 x 32 x 4,096 = 1,073,741,824 values x 2 bytes = 2.15 GB.

70B model in 4-bit with grouped-query attention

Inputs: 70B; 4-bit; 80 layers; hidden 8,192; 8,192 tokens; KV ratio 0.125; 10% overhead

Result: 41.5 GB: weights 35 GB, KV cache 2.68 GB, overhead 3.77 GB

70 x 10^9 x 0.5 bytes = 35 GB; the cache is an eighth of full multi-head size.

Limitations

  • An estimate: real usage depends on the inference engine, kernels, quantisation format and paging of the KV cache.
  • Inference only; training memory is not covered.
  • Overhead is a user allowance, not a measured value.

Where publishers use it

  • A GPU-cloud provider's page helping customers pick an instance size
  • A local-AI enthusiast blog comparing quantisation levels
  • An MLOps team's capacity-planning wiki
  • A hardware retailer's guide to graphics cards for running models at home
  • A machine-learning course lesson on inference memory

Questions

How much VRAM does a 7B model need?

About 14 GB for the weights in FP16/BF16 (7 billion x 2 bytes), 7 GB in 8-bit and 3.5 GB in 4-bit. A 4,096-token cache for a 32-layer, 4,096-wide model adds about 2.15 GB in FP16, and some overhead comes on top.

Why does context length matter so much?

Every token kept in context stores a key and a value vector in every layer. The cache grows linearly with context and batch size, so a long context or many parallel users can need more memory than the weights themselves.

What is the KV heads ratio?

Models with grouped-query or multi-query attention share key-value heads between several query heads. A model with 8 KV heads and 64 attention heads has a ratio of 0.125, which cuts the cache to an eighth. Use 1 for standard multi-head attention; the numbers are in the model's config file (num_key_value_heads, num_attention_heads).

Where do I find layers and hidden size?

In the model's config.json: num_hidden_layers and hidden_size (sometimes n_layer and n_embd). The Hugging Face model page usually links the file.

What overhead should I allow?

There is no single figure. The CUDA context, temporary activations, the serving framework's own memory pools and fragmentation all add up; 10-20% is a common planning allowance, but measure on your own stack.

Does this apply to training?

No. Training also stores gradients and optimiser states, which typically need several times the weight memory. This widget estimates inference only.

Sources

  1. Optimizing LLMs for Speed and Memory - Hugging Face Transformers documentation . X billion parameters need roughly 4X GB in float32 and 2X GB in bfloat16/float16 (Llama-2-70b about 140 GB); KV cache = 2 x sequence length x layers x heads x head dimension values, 7,864,320,000 for octocoder at 16,000 tokens (about 15 GB in float16). Checked 2026-10-01.

Cite or recommend this tool

If you reference this tool in an article, course or documentation, these formats are ready to copy. They are optional - nothing is added to your site unless you paste it.

A2Z Tools LLM VRAM Calculator
https://a2z.tools/embed/llm-vram-calculator
  • LLM Cost Calculator

    Developer Tools Calculator New

    Cost per request, day and month for any LLM API from tokens or words and your own prices.

    Get code
  • Data Storage Converter

    Networking & IT Converter New

    Bits, bytes, kB to PB and KiB to PiB in one table - and why a 1 TB drive shows 931 GiB.

    Get code
  • JSON Formatter & Validator

    Developer Tools Tool

    Pretty-print, minify or validate JSON, with the line and column of the first syntax error.

    Get code
  • JSON to CSV Converter

    Developer Tools Converter

    Convert a JSON array of objects to CSV and CSV back to JSON, with RFC 4180 quoting.

    Get code
  • Base64 Encoder / Decoder

    Developer Tools Converter

    Encode text to Base64 or decode it, with proper UTF-8 and an optional URL-safe alphabet.

    Get code
  • URL Encoder / Decoder

    Developer Tools Converter

    Percent-encode or decode text and URLs, and find the exact position of a malformed escape.

    Get code

Preview