🛡️

Security & Guardrails

AI safety evaluation, red-teaming, LLM guardrails, vulnerability scanning, and compliance audit tools

67 projects

Anthropic Cybersecurity Skills

32.0k · Python
Active A+

754 structured cybersecurity skills for AI agents mapped to 5 frameworks: MITRE ATT&CK, NIST CSF 2.0, MITRE ATLAS, D3FEND and NIST AI RMF. Works with Claude Code, Codex CLI, Cursor, Gemini CLI and 20+ platforms.

pythonsecurityagent +2
  • · 754 structured cybersecurity skills spanning 26 security domains in agentskills.io format
  • · Unified cross-framework mapping to MITRE ATT&CK v19.1, NIST CSF 2.0, MITRE ATLAS, D3FEND, and NIST AI RMF
  • · Compatible with Claude Code, GitHub Copilot, Codex CLI, Cursor, Gemini CLI and 20+ AI platforms

Promptfoo

24.8k · TypeScript
Active A

Test and evaluate LLM prompts, agents, and RAG pipelines. Built-in red teaming and security evaluation for reliable AI applications.

testingevaluationred-teaming +2
  • · Automated LLM evaluations — Batch test prompts, models, and RAG pipeline output quality
  • · Red team security testing — Built-in vulnerability scanning and adversarial testing for LLM security
  • · Multi-model comparison — Side-by-side comparison of OpenAI, Anthropic, Azure, Bedrock, Ollama models

PentAGI

22.3k · Go
Active A+

Fully autonomous AI Agents system capable of performing complex penetration testing tasks using multi-agent architecture with support for multiple LLM providers.

securitytestingmulti-agent +2
  • · Fully autonomous AI agent for penetration testing with intelligent task planning
  • · Sandboxed Docker environment ensuring complete isolation of all operations
  • · 20+ built-in professional security tools including nmap, metasploit, and sqlmap

SWE-agent

20.2k · Python
Active A+

SWE-agent takes a GitHub issue and automatically generates fixes using your LLM of choice, also applicable to cybersecurity auditing and competitive coding. NeurIPS 2024 paper.

swecodingagent +2
  • · State-of-the-art performance on SWE-bench among open-source software engineering agents
  • · Supports multiple LLM backends including GPT-4o and Claude Sonnet 4
  • · Configurable via a single YAML file with full documentation

OpenAI Evals

19.4k · Python
Stale B

OpenAI's framework for evaluating LLMs and LLM systems, providing an open-source registry of benchmarks and tools for systematic model assessment.

llm-evaluationbenchmarkevals +2
  • · Open-source registry of evals for testing different dimensions of LLM performance
  • · Custom eval creation using basic and model-graded templates without writing code
  • · Private evals support for evaluating LLM patterns in your workflow without exposing data

SkillSpector

15.7k · Python
Active A

NVIDIA's SkillSpector inspects and evaluates the tool-use and function-calling skills of LLM agents against safety, correctness, and performance criteria.

security-guardrailsmcpstatic-analysis +1
  • · Open-sourced by NVIDIA with backing from a hardware leader
  • · Static scanning of agent skills to detect malicious code and vulnerabilities
  • · Multi-dimensional detection: command injection, credential leaks, unsafe IO

PentestGPT

15.2k · Python
Normal A

An automated penetration testing agentic framework powered by large language models for security testing and vulnerability discovery.

penetration-testingsecurityllm +2
  • · AI-powered autonomous penetration testing agent with agentic pipeline for intelligent security assessment
  • · Docker-first isolated environment with 100+ security tools pre-installed and reproducible setup
  • · Session persistence to save and resume penetration testing sessions across multiple runs

OpenSandbox

14.9k · Go
Active A

OpenSandbox is an open-source, secure, fast, and extensible sandbox runtime for AI agents, developed by Alibaba.

sandboxai-infrastructurekubernetes +2
  • · Multi-language SDKs: Python, Java/Kotlin, JavaScript/TypeScript, C#/.NET, and Go with unified sandbox APIs
  • · Docker and Kubernetes runtimes: built-in lifecycle management for both local development and large-scale distributed scheduling
  • · Strong isolation: supports gVisor, Kata Containers, and Firecracker microVM secure container runtimes

E2B

13.7k · Python
Active A+

E2B provides secure cloud sandboxes for AI agents, supporting code execution, file operations, and isolated compute as an execution layer for coding and automation workflows.

sandboxcode-executionsecurity +1
  • · Secure Cloud Sandboxes: Provides isolated cloud execution environment for AI agents
  • · Code Execution: Supports JavaScript/Python SDK to run AI-generated code
  • · File Operations: Perform file read/write and management within sandboxes

Portkey AI Gateway

12.9k · TypeScript
Stale B

Portkey AI Gateway is a blazing fast AI gateway with integrated guardrails, routing to 200+ LLMs with 50+ AI guardrails through a single fast and friendly API.

gatewayllm-routingguardrails +2
  • · Route to 250+ LLMs with a single OpenAI-compatible API endpoint
  • · Blazing fast at under 1ms added latency with a 122kb footprint
  • · Built-in automatic retries, fallbacks, and load balancing for reliability

HexStrike AI

11.5k · Python
Active A

HexStrike AI is an advanced MCP server that lets AI agents autonomously run 150+ cybersecurity tools for automated pentesting, vulnerability discovery, and security research.

cybersecuritypentestingmcp-server +2
  • · 150+ professional security tools integrated — network recon, web app testing, password cracking, binary analysis, and cloud security
  • · 12+ autonomous AI agents for specialized tasks like BugBounty, CTF solving, CVE intelligence, and exploit generation
  • · MCP protocol integration with Claude, GPT, VS Code Copilot, Cursor, and other MCP-compatible clients

Presidio

10.7k · Python
Active A+

Microsoft's open-source context-aware PII detection and de-identification SDK for text, images, and structured data, providing sensitive data protection for LLM applications and agents.

pii-detectiondata-maskingprivacy +2
  • · Context-aware PII detection — Identifies credit card numbers, names, addresses, and other sensitive entities using NER, regex, rule logic, and checksums
  • · Multiple de-identification modes — Supports masking, replacement, encryption, pseudonymization, and other anonymization strategies
  • · Image PII redaction — Built-in image text recognition and PII region masking, with DICOM medical image support

GhidraMCP

9.9k · Java
Stale B

MCP server for Ghidra reverse engineering platform, enabling AI agents to autonomously perform binary analysis and vulnerability discovery.

mcpreverse-engineeringghidra +2
  • · MCP server for Ghidra enabling LLMs to autonomously perform binary reverse engineering
  • · Decompile and analyze binaries with automatic renaming of methods and data
  • · List all methods, classes, imports, and exports, exposing core Ghidra functionality to MCP clients

CAI

9.8k · Python
Active A

Alias Robotics' open-source AI security research agent framework for multi-agent orchestration of cybersecurity tasks, integrating 300+ AI models, designed for red-team operations and security research.

cybersecurityai-agentsred-team +2
  • · Multi-agent orchestration — Built-in cybersecurity-specialized agent roles for reconnaissance, vulnerability scanning, exploitation, and reporting
  • · 300+ model support — Unified interface to call OpenAI, Anthropic, Gemini, Ollama, and other major LLM providers
  • · Penetration testing tool integration — Built-in calls for common pentest tools, supports CTF benchmarks

Garak

9.1k · Python
Active A

NVIDIA's open-source LLM vulnerability scanner that automatically detects security issues in language models including safety vulnerabilities, hallucination tendencies, jailbreak risks, and prompt injection attacks.

llm-securityvulnerability-scannerllm-evaluation +2
  • · Automated LLM vulnerability scanning for hallucination, data leakage, prompt injection, and jailbreak detection
  • · Static, dynamic, and adaptive probe combinations for comprehensive security assessment
  • · Broad LLM support: Hugging Face, OpenAI, AWS Bedrock, Replicate, llama.cpp, and REST-accessible models

OpenShell

8.5k · Rust
Active A

OpenShell is the safe, private runtime for autonomous AI agents, developed by NVIDIA. Provides controlled execution environments and resource management.

rustagentframework +2
  • · Sandboxed execution environments protecting data, credentials, and infrastructure with declarative YAML policies
  • · Four defense-in-depth layers: filesystem, network, process, and inference policy enforcement
  • · Hot-reloadable network and inference policies via `openshell policy set` without restarting sandboxes

Guardrails AI

7.3k · Python
Active A+

Guardrails AI adds programmable guardrails to large language models, ensuring reliability and safety through input/output validation, structured data extraction, and custom validators.

guardrailsllm-safetyvalidation +2
  • · Input/Output Guards — detect, quantify, and mitigate specific risk types in LLM applications
  • · Guardrails Hub pre-built validators — RegexMatch, CompetitorCheck, ToxicLanguage and more out of the box
  • · Pydantic structured output — enforce LLM output format via function calling or prompt optimization

Agent Governance Toolkit

6.2k · Python
Active A

Microsoft's AI Agent Governance Toolkit providing policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.

securityevaluationpython +2
  • · Deterministic policy enforcement engine: intercepts and evaluates all operations before model output reaches tool calls, with YAML policy files defining allow/deny/approval-required rules — denied actions are structurally impossible
  • · Zero-trust identity and audit trail: assigns unique DIDs to each agent, generates tamper-evident audit records for every operation including active policy, request content, and decision rationale
  • · Multi-language SDK support: Python, TypeScript, .NET, Rust, and Go SDKs with unified policy evaluation APIs covering major development stacks

AI-Infra-Guard

6.1k · Python
Active A+

Tencent's full-stack AI red teaming platform integrating OpenClaw security scanning, agent scanning, skills scanning, MCP scanning, AI infrastructure scanning, and LLM jailbreak evaluation.

ai-securityred-teamingllm-security +2
  • · ClawScan: OpenClaw security scanning for detecting vulnerabilities in OpenClaw configurations and skills
  • · Agent Scan: vulnerability scanning across AI agents with WebSocket provider support
  • · AI infrastructure vulnerability scan covering 68+ AI components with 600+ CVE rules

(24 / 67)

Related Articles

Agent 评估LLM 评测自动化测试

Agent Evaluation and Testing: From Vibe Checks to End-to-End Pipelines

Most teams evaluate agents by checking a few examples. Real evaluation needs layered metrics, non-rotting datasets, and judges that push back. This article provides runnable code patterns and a practical decision framework.

RAGhallucination-detectionagent-evaluation

Agent Hallucination Defense: Practical Mitigation Patterns Beyond Guardrails

Why do LLM agents hallucinate? This article traces root causes and systematically reviews practical mitigation patterns: retrieval augmentation, confidence scoring, multi-agent cross-validation, forced citation backtracking, and observability with UpTrain, Giskard, RagaAI Catalyst, Comet Opik, and NVIDIA Garak.

安全Prompt InjectionOWASP

Agent Prompt Injection Defense: OWASP LLM01 in Practice

Based on OWASP LLM Top 10 engineering practice, this article systematically explains the seven layers of defense-in-depth for agent prompt injection: input sanitization, instruction isolation, least-privilege, output auditing, guardrails frameworks, continuous red-teaming, and kill switches -- with actionable code and toolchains.

sandbox-executionmicrovmfirecracker

Sandboxing Code Execution in AI Agents: From Docker to microVMs, a Decision Matrix

A side-by-side comparison of five sandbox technologies, weighing latency, security, and ops cost.

security-guardrailsred-teamprompt-injection

AI Agent Guardrails and Red Teaming in Practice: From Rule Engines to Adversarial Evaluation

Five-layer defense plus red-team loop, built on five open-source projects you can copy.

AI Agent安全Prompt Injection

AI Agent Security in Practice: From Prompt Injection to Defense in Depth

A systematic walkthrough of three major attack surfaces in AI agents, with practical code examples for prompt injection defense, tool permission scoping, and output filtering.