mradermacher/Gamunu-4B-Instruct-Alpha-i1-GGUF
mradermacher/Gamunu-4B-Instruct-Alpha-i1-GGUF is a question answering model indexed for deployment research. Estimated minimum GPU memory is 6 GB. It is publicly listed on Hugging Face.
What to verify before deployment
- Confirm the exact weight format and quantization.
- Measure memory at your intended context or image size.
- Review model-card limitations and evaluation methodology.
- Test latency and throughput on your target runtime.
Likely commercial-friendly
License metadata is an index signal, not legal advice. Follow the repository license and any model-specific acceptable-use terms.
Verify at the source →More question-answering models
mradermacher/Gamunu-4B-Instruct-Alpha-i1-GGUF
mradermacher/Gamunu-4B-Instruct-Alpha-i1-GGUF is a question answering model indexed for deployment research. Estimated minimum GPU memory is 6 GB. It is publicly listed on Hugging Face.
prithivMLmods/Code-as-World-VL-4B-GGUF
prithivMLmods/Code-as-World-VL-4B-GGUF is a question answering model indexed for deployment research. Estimated minimum GPU memory is 6 GB. It is publicly listed on Hugging Face.
mradermacher/medgemma-4b-pt-CPT-SFT-i1-GGUF
mradermacher/medgemma-4b-pt-CPT-SFT-i1-GGUF is a question answering model indexed for deployment research. Estimated minimum GPU memory is 6 GB. It is publicly listed on Hugging Face.
X054848/Apollo2-0.5B-Q8_0-GGUF
X054848/Apollo2-0.5B-Q8_0-GGUF is a question answering model indexed for deployment research. Estimated minimum GPU memory is 4 GB. It is publicly listed on Hugging Face.
Satya-Dey/qwen2.5-1.5b-medmcqa-qlora-v2
Satya-Dey/qwen2.5-1.5b-medmcqa-qlora-v2 is a question answering model indexed for deployment research. Estimated minimum GPU memory is 6 GB. It is publicly listed on Hugging Face.
maianh511/internvl2_1b_finetune_lora_viet_chart_vqa
maianh511/internvl2_1b_finetune_lora_viet_chart_vqa is a question answering model indexed for deployment research. Estimated minimum GPU memory is 6 GB. It is publicly listed on Hugging Face.
Operate a GPU cloud or inference API?
Reach developers after they have selected a model and are ready to run it.