groxaxo/Nemotron-3-Embed-1B-AWQ-W4A16
groxaxo/Nemotron-3-Embed-1B-AWQ-W4A16 is a feature extraction model indexed for deployment research. Estimated minimum GPU memory is 6 GB. It is publicly listed on Hugging Face.
What to verify before deployment
- Confirm the exact weight format and quantization.
- Measure memory at your intended context or image size.
- Review model-card limitations and evaluation methodology.
- Test latency and throughput on your target runtime.
Human review required
License metadata is an index signal, not legal advice. Follow the repository license and any model-specific acceptable-use terms.
Verify at the source →More feature-extraction models
BAAI/bge-small-en-v1.5
BAAI/bge-small-en-v1.5 is a feature extraction model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.
BAAI/bge-large-en-v1.5
BAAI/bge-large-en-v1.5 is a feature extraction model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.
BAAI/bge-base-en-v1.5
BAAI/bge-base-en-v1.5 is a feature extraction model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.
Qwen/Qwen3-Embedding-0.6B
Qwen/Qwen3-Embedding-0.6B is a feature extraction model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.
intfloat/multilingual-e5-large
intfloat/multilingual-e5-large is a feature extraction model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.
ibm-granite/granite-embedding-small-english-r2
ibm-granite/granite-embedding-small-english-r2 is a feature extraction model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.
Operate a GPU cloud or inference API?
Reach developers after they have selected a model and are ready to run it.