Loading…
Democratizing AI Safety with RiskRubric.ai
Hugging FaceGal Moyal
Summary
Cloud Security Alliance and Noma Security introduced RiskRubric.ai to provide standardized, transparent risk assessments across the open AI model ecosystem. The framework evaluates AI models across six pillars—transparency, reliability, security, privacy, safety, and reputation—using over 1,000 reliability tests, 200 adversarial security probes, automated code scanning, and harmful content evaluations. Each model receives 0–100 scores and A–F letter grades, supplemented by specific vulnerability findings and recommended mitigation strategies to assist deployment filtering. Initial benchmark results across models showed composite scores ranging from 47 to 94 with a median of 81, revealing polarized safety distributions and indicating that security hardening directly correlates with reduced safety risks.
Context
With more than 500,000 models available on the Hugging Face hub, developers lack a systematic, standardized method to evaluate model security posture, privacy implications, and potential failure modes prior to deployment.
Approach / What changed
Cloud Security Alliance and Noma Security launched RiskRubric.ai to evaluate models across six pillars: transparency, reliability, security, privacy, safety, and reputation. The platform automates evaluations using over 1,000 reliability tests, 200 adversarial security probes for jailbreaks and prompt injections, automated code scanning, documentation reviews, privacy assessments, and structured harmful content tests. These produce 0-100 scores and A-F letter grades alongside vulnerability reports and remediation guidance.
Takeaways
- Model risk scores ranged from 47 to 94 with a median of 81, showing polarization where 54 percent reached A or B levels while a long tail clustered in the medium-to-low protection C and D range.
- Safety pillar scores exhibited the widest variation across evaluated models but tracked closely with security posture, indicating that prompt injection defenses and policy enforcement directly mitigate harmful outputs.
- Stricter guardrails frequently decrease user-perceived transparency through opaque refusals, which can be mitigated by combining safeguards with explanatory refusals and provenance signals.
Related reading
Amazon ·
Building trust into AI
Amazon integrates responsible AI practices across four core model development phases: pretraining, post-training, evaluation, and third-party monitoring. During pretraining, researchers augment training corpuses with specialized safety datasets, multimodality alignments, and learning exercises to teach foundational safety concepts instead of purely filtering harmful text. In post-training, reinforcement learning from human feedback uses human preference rankings, auxiliary-reward models, and independent LLM judges to ensure model outputs adhere to safety policies. For specialized use cases requiring access to sensitive domains like security testing, researchers apply low-rank adaptors to alter model behaviors surgically without retraining base weights. Cross-functional policy teams guide these technical stages by mapping risks against eight responsible AI dimensions and continuously updating behavioral boundaries to reflect evolving regulatory frameworks.
Staff writerMeta ·
How We’re Building Scam Alert on WhatsApp With End-to-End Encryption and Verifiability Guarantees