← All guides
8GB GPU MODEL GUIDE

Best Hugging Face models for an 8GB GPU

An 8GB graphics card can run useful open models, but the repository name alone does not tell you whether the selected weights, context and runtime will fit.

Quick answer

Start with smaller language, embedding, classification or compact image models. For generative LLMs, 4-bit quantization is often necessary once parameter counts move beyond the smallest families. Leave memory for the runtime, KV cache and input—not just weights.

Live 8GB shortlist

The models below are currently estimated by AI Pentium to require no more than 8GB of minimum VRAM. They are ordered by Hugging Face download metadata and update automatically with the catalog.

sentence-similarity

sentence-transformers/all-MiniLM-L6-v2

sentence-transformers/all-MiniLM-L6-v2 is a sentence similarity model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.

253,331,9945,596
Unknown4 GB+ VRAMapache-2.0
Deployment details
text-ranking

cross-encoder/ms-marco-MiniLM-L6-v2

cross-encoder/ms-marco-MiniLM-L6-v2 is a text ranking model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

86,661,012314
Unknown8 GB+ VRAMapache-2.0
Deployment details
feature-extraction

BAAI/bge-small-en-v1.5

BAAI/bge-small-en-v1.5 is a feature extraction model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.

64,933,788550
Unknown4 GB+ VRAMmit
Deployment details
general AI

google/electra-base-discriminator

google/electra-base-discriminator is a AI workload model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

58,160,168158
Unknown8 GB+ VRAMapache-2.0
Deployment details
fill-mask

google-bert/bert-base-uncased

google-bert/bert-base-uncased is a fill mask model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

50,396,5173,008
Unknown8 GB+ VRAMapache-2.0
Deployment details
sentence-similarity

BAAI/bge-m3

BAAI/bge-m3 is a sentence similarity model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.

37,985,5663,483
Unknown4 GB+ VRAMmit
Deployment details
sentence-similarity

sentence-transformers/all-mpnet-base-v2

sentence-transformers/all-mpnet-base-v2 is a sentence similarity model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.

23,908,2571,351
Unknown4 GB+ VRAMapache-2.0
Deployment details
time-series-forecasting

amazon/chronos-2

amazon/chronos-2 is a time series forecasting model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

23,768,593435
Unknown8 GB+ VRAMapache-2.0
Deployment details
translation

google-t5/t5-small

google-t5/t5-small is a translation model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

23,696,426600
Unknown8 GB+ VRAMapache-2.0
Deployment details
text-generation

Qwen/Qwen3-0.6B

Qwen/Qwen3-0.6B is a text generation model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.

21,493,0271,588
600M4 GB+ VRAMapache-2.0
Deployment details
fill-mask

FacebookAI/xlm-roberta-base

FacebookAI/xlm-roberta-base is a fill mask model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

21,410,310895
Unknown8 GB+ VRAMmit
Deployment details
zero-shot-image-classification

openai/clip-vit-base-patch32

openai/clip-vit-base-patch32 is a zero shot image classification model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

20,702,7631,225
Unknown8 GB+ VRAMLicense unknown
Deployment details
general AI

Comfy-Org/MiniMax-H3

Comfy-Org/MiniMax-H3 is a AI workload model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

18,804,1521,730
Unknown8 GB+ VRAMother
Deployment details
text-classification

BAAI/bge-reranker-v2-m3

BAAI/bge-reranker-v2-m3 is a text classification model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

18,159,9061,162
Unknown8 GB+ VRAMapache-2.0
Deployment details
general AI

Comfy-Org/stable-diffusion-v1-5-archive

Comfy-Org/stable-diffusion-v1-5-archive is a AI workload model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

17,283,732127
Unknown8 GB+ VRAMcreativeml-openrail-m
Deployment details
image-classification

timm/mobilenetv3_small_100.lamb_in1k

timm/mobilenetv3_small_100.lamb_in1k is a image classification model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

16,248,411103
Unknown8 GB+ VRAMapache-2.0
Deployment details
sentence-similarity

nomic-ai/nomic-embed-text-v1.5

nomic-ai/nomic-embed-text-v1.5 is a sentence similarity model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.

16,210,409903
Unknown4 GB+ VRAMapache-2.0
Deployment details
text-generation

openai-community/gpt2

openai-community/gpt2 is a text generation model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

14,748,3563,732
Unknown8 GB+ VRAMmit
Deployment details
automatic-speech-recognition

jonatasgrosman/wav2vec2-large-xlsr-53-japanese

jonatasgrosman/wav2vec2-large-xlsr-53-japanese is a automatic speech recognition model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

13,506,34163
Unknown8 GB+ VRAMapache-2.0
Deployment details
feature-extraction

BAAI/bge-large-en-v1.5

BAAI/bge-large-en-v1.5 is a feature extraction model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.

13,169,864721
Unknown4 GB+ VRAMmit
Deployment details
sentence-similarity

intfloat/multilingual-e5-small

intfloat/multilingual-e5-small is a sentence similarity model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.

12,313,873399
Unknown4 GB+ VRAMmit
Deployment details
image-classification

timm/efficientnet_b3.ra2_in1k

timm/efficientnet_b3.ra2_in1k is a image classification model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.

12,216,3187
Unknown8 GB+ VRAMapache-2.0
Deployment details

What “fits in 8GB” really means

How to verify before downloading

  1. Open the model detail page and note the parameter count and estimated minimum VRAM.
  2. Follow the source link to inspect the exact files and supported quantizations.
  3. Begin with batch size one and a conservative context or image size.
  4. Measure peak allocated memory, then increase the workload gradually.

This list is a discovery aid, not a benchmark or guarantee. Always test the exact Hugging Face repository and runtime.

Open the live 8GB hardware matcher