Lv63
2750 积分 2024-10-23 加入
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
12天前
已完结
SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions
12天前
已完结
Scalable Extraction of Training Data from (Production) Language Models
12天前
已完结
Extracting Training Data from Large Language Models
12天前
已完结
AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents
12天前
已完结
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
12天前
已完结
SecAlign: Defending Against Prompt Injection with Preference Optimization
12天前
已完结
Ignore Previous Prompt: Attack Techniques For Language Models
12天前
已完结
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
12天前
已完结
Jailbreaking Black Box Large Language Models in Twenty Queries
12天前
已完结