moderation-prompts
updated
mmathys/openai-moderation-api-evaluation
Viewer
• Updated • 1.68k • 1.05k
• 38
Viewer
• Updated • 169k • 34k
• 1.95k
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks,
and Refusals of LLMs
Paper
• 2406.18495
• Published • 15
ShieldGemma: Generative AI Content Moderation Based on Gemma
Paper
• 2407.21772
• Published • 15
Viewer
• Updated • 1M • 7.03k
• 965
PKU-Alignment/BeaverTails
Viewer
• Updated • 364k • 23.4k
• 111
AgentPublic/camembert-base-toxic-fr-user-prompts
Text Classification
• 0.1B • Updated • 30
• 8
Viewer
• Updated • 30.7k • 1.92k
• 31
meta-llama/Llama-Guard-3-8B
Text Generation
• 8B • Updated • 165k
• • 313
davanstrien/aart-ai-safety-dataset
Viewer
• Updated • 3.27k • 16
• 2
Viewer
• Updated • 520 • 16.2k
• 110