AI paper index
RefusalGuard-M: a scalable human–machine framework for multi-turn LLM jailbreak evaluation via semantic refusal manifold modeling
One-line summary
An AI research paper on RefusalGuard-M: a scalable human–machine framework for multi-turn LLM jailbreak evaluation via semantic refusal manifold modeling.
Engineering notes
Engineering notes will be added by the aipentium editorial team.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。
Original abstract
Abstract Existing multi-turn jailbreak evaluation methods increasingly rely on large language models (LLMs) as automated judges to reduce the cost and scalability limitations of human assessment. However, recent studies show that LLM-based evaluators can diverge from human judgments under adversarial strategies involving subtle linguistic and semantic variations, raising reliability concerns in safety-critical domains such as cybersecurity. To address this challenge, we propose Refusal Manifold Guard (RefusalGuard-M), an open-source semantic evaluation framework that constructs a semantic refusal manifold from human-validated refusal responses for assessing LLM jailbreak interactions, including multi-turn scenarios. RefusalGuard-M uses embedding-based geometric representations to measure deviations from refusal behavior, providing a lightweight, interpretable, and reproducible alternative to LLM-based judging. We evaluate the framework across AdvBench, HarmBench, and CyMulTenSet, covering diverse jailbreak strategies, linguistic transformations, and multi-turn scenarios. Results show that RefusalGuard-M achieves strong agreement with human annotations and comparable recall performance to GPT-based evaluators while adopting a conservative evaluation strategy that prioritizes the detection of harmful outputs. On CyMulTenSet, which evaluates past-tense reformulated multi-turn jailbreaks, RefusalGuard-M achieves up to 0.87 recall, compared with 0.86 for GPT-5 and 0.81 for GPT-4, and reduces inference overhead by up to 3.7 $$\times$$ × relative to embedding-based baselines. These findings demonstrate that semantic refusal representations provide an efficient and scalable approach for jailbreak evaluation, particularly in cybersecurity settings where minimizing missed harmful outputs is critical.
Links and sources
Need this topic turned into a technical roadmap?
aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.
Request B2B AI research
Comments