jinaai/jina-reranker-v3.5-GGUF
jinaai/jina-reranker-v3.5-GGUF is a text ranking model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.
What to verify before deployment
- Confirm the exact weight format and quantization.
- Measure memory at your intended context or image size.
- Review model-card limitations and evaluation methodology.
- Test latency and throughput on your target runtime.
Human review required
License metadata is an index signal, not legal advice. Follow the repository license and any model-specific acceptable-use terms.
Verify at the source →More text-ranking models
cross-encoder/ms-marco-MiniLM-L6-v2
cross-encoder/ms-marco-MiniLM-L6-v2 is a text ranking model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.
cross-encoder/ms-marco-MiniLM-L4-v2
cross-encoder/ms-marco-MiniLM-L4-v2 is a text ranking model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.
Alibaba-NLP/gte-reranker-modernbert-base
Alibaba-NLP/gte-reranker-modernbert-base is a text ranking model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.
cross-encoder/mmarco-mMiniLMv2-L12-H384-v1
cross-encoder/mmarco-mMiniLMv2-L12-H384-v1 is a text ranking model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.
Qwen/Qwen3-VL-Reranker-2B
Qwen/Qwen3-VL-Reranker-2B is a text ranking model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.
cross-encoder/ms-marco-MiniLM-L12-v2
cross-encoder/ms-marco-MiniLM-L12-v2 is a text ranking model indexed for deployment research. Estimated minimum GPU memory is 8 GB. It is publicly listed on Hugging Face.
Operate a GPU cloud or inference API?
Reach developers after they have selected a model and are ready to run it.