Proof Benchmark - Adversarial AI Security

The benchmark that
produces proof,
not evidence.

Formal verification proves what a system can do. Proof of Control proves what it did. Proof of Efficacy™ proves what survives when someone tries to break it. AI security has dozens of frameworks covering the first two. Almost nothing covers the third. Every other benchmark produces a document. Proof Benchmark produces a tamper-resistant, independently witnessed, publicly verifiable execution record - anchored to the NIST Randomness Beacon before a single test case runs. The result is either a ProofStamp or it isn't. There is no middle tier.

Registry Status - Live
Threat categories covered
30+
Certified runs on record
1
Reference run - containment
99.2%
Reference run - detection
100%
Applicable cases (ref run)
164
Frameworks replaced
0
Frameworks made provable
30+
Proof of Efficacy™
ProofRegister - Products that have run
The round trip has to be adversarial: a hostile action executes against the target, and the response proves whether it was stopped. A transaction that cannot fail is not proof - it is just a logged transaction. Proof has to stand the test of failure. That is why runs anchor to the NIST Randomness Beacon before execution, not after: a tamper-evident log only proves a result wasn't altered once recorded. Pre-commitment proves the outcome wasn't chosen before it was recorded - the difference that matters when the claim is about failure being genuinely possible.
Class A
≥ 90%
Stamped
Class B
≥ 80%
Stamped
Class C
≥ 70%
Stamped
Class D
≥ 60%
No stamp
Class E
≥ 50%
No stamp
Class F
< 50%
No stamp
Class is determined by Containment (PES - Proof of Efficacy Score). Every run publishes to ProofRegister regardless of class. A PROOFstamp™ is issued only at Class C or above.
Pipelock v3.0.0
Containment (PES)
99.2%
Detection Rate
100%
False Positive Rate
4.5%
Applicable Cases
164
Proof Record: PR-2026-00028  ·  NIST Beacon Block: 29  ·  Pulse Index: 1852788
Corpus: Proof Benchmark: Agentic AI v1.0 - The Gauntlet
117 per-TTP child proof records (proofType 2) · Bitcoin OP_RETURN anchored · NIST Beacon sealed
Administered by: HACKERverse® Independent Test Lab (non-voting, recused from governance)
Pre-Dated
PROOFstamp™
Class A · Executing ✓

proofstamp.io
CHOMP AND SHRED - Ghost Corpus™ · TrendAI™ EDR/XDR
Built and delivered for TrendAI™: a working, dynamically changing corpus - not a fixed set, a body of test material that updates as new campaigns are found in the wild
Corpus: CHOMP AND SHRED - Ghost Corpus™ - campaigns sourced live, converted to executable test in minutes · NIST Beacon sealed
CHOMP Engine: Continuous Campaign Injection - ran against TrendAI's™ target environment until TrendAI shut down that lab. Paused pending a replacement target
Administered by: HACKERverse® - Internal (no independent test lab currently exists for this bench; HACKERverse tested on TrendAI's™ behalf pending PESA lab accreditation)
Paused - Seeking Target
PROOFstamp™
Not Yet Issued

proofstamp.io
ProofCorpus™ - Super Corpus
BountyBench
46
CVE-Bench
40
SEC-Bench
200
CyberGym (pulling)
1,507
Merged corpus: 286 entries (growing toward full set as CyberGym completes)  ·  Repo: proofprotocol/proofcorpus (private, pending license review)
A merged, provenance-tracked vulnerability test corpus unifying four open academic benchmarks into one attributable source - the raw material future Proof of Efficacy™ benches draw from
Administered by: HACKERverse® - Internal (no independent test lab currently exists for this corpus; HACKERverse is assembling and maintaining ProofCorpus internally pending PESA lab accreditation)
Building
PROOFstamp™
Not Yet Issued

proofstamp.io
Threat Corpus - Living Document · v1.0 · 2026-07-14 · 30+ categories
Cloud and Infrastructure
  • Cloud Misconfiguration Exploitation
  • Container Escape
  • Kubernetes Cluster Compromise
  • Serverless and Function Abuse
  • CI/CD Pipeline Compromise
  • API Abuse and Business Logic
AI and Agentic
  • AI Agent Impersonation
  • Prompt Injection - Direct and Indirect
  • Multi-Agent Collusion
  • Agentic Supply Chain Compromise
  • Model Poisoning and Training Data Attacks
  • AI-Generated Malware and Synthetic TTPs
  • Adversarial AI vs Defensive AI
  • Agent Authorization Escalation
  • Agentic Worms and Self-Propagating Agents
  • Synthetic Identity and Deepfake Social Engineering
  • AI Hallucination Exploitation
  • Agent Memory and Context Poisoning
Future and Day-Zero
  • Pre-CVE Zero-Day Exploitation
  • Novel Campaign Stream Detection
  • AI-Synthesized Zero-Day Techniques
  • Quantum-Enabled Cryptographic Attacks
  • Emergent Multi-Agent Threat Behaviors
  • Infrastructure AI Takeover
  • Agentic Ransomware
  • Regulatory Evidentiary Mandate
Verticals
  • Financial Services
  • Healthcare
  • Critical Infrastructure (OT/ICS/SCADA)
  • Defense and Government
  • Telecommunications
Proof of Control™
Frameworks, controls, and taxonomies referenced across the industry
Signed requests and responses on a real transaction are attestation, not proof. The action had nothing to fail against. Proof of Control is what these frameworks provide even at their most rigorous.
Category What it answers
ThreatsWhat to worry about
VulnerabilitiesWhat can go wrong
PrinciplesWhat good looks like
ControlsWhat to enforce
GovernanceWhat to build
Runtime enforcementWhat to enforce at the action boundary
Evidence and verificationHow you show any of the above actually happened
Adversarial efficacyWhether it held when someone tried to break it
Framework Owner What It Does
MAESTROCSA / Ken Huang (Feb 2025)7-layer threat identification for agentic AI
AICMCSA243 controls across 18 domains for AI systems
ATFCSAOperationalizes AICM for agents; Zero Trust framing
STAR for AICSAI FoundationCatastrophic risk annex to CSA STAR; phases through Dec 2027
OWASP LLM Top 10OWASP (Aug 2023, updated 2025)Ranked list of LLM-specific vulnerabilities
OWASP Agentic Threats v1.0aOWASP (Feb 2025)15 agentic threat categories
OWASP Multi-Agentic Threat Modeling GuideOWASP (Apr 2025)Structured threat modeling for multi-agent systems
OWASP Top 10 for Agentic Apps 2026OWASP (Dec 2025)Top 10 agentic application risks, 100+ contributors
MITRE ATLASMITREAdversarial ML technique catalog, living knowledge base
NIST AI RMF 1.0NISTGovern, map, measure, manage AI risk
NIST AI 600-1NISTRisk management applied to generative AI
NIST AI 100-2NISTAdversarial ML attacks and mitigations terminology
NIST CAISI Agent Standards InitiativeNIST (Feb 2026)US government AI agent standards program
COSAiSNISTSP 800-53 control overlays for five AI use cases
Google SAIF / SAIF 2.0GoogleEmbed security into AI model development; Agent Risk Map
Cisco DefenseClawCisco (Mar 2026)Skills scanner, MCP scanner, AI BOM, sandboxing
Cisco AI DefenseCiscoRuntime guardrails and red teaming for agentic workflows
Palo Alto AIRS 3Palo Alto NetworksAgent lifecycle security
CrowdStrike AI Runtime ProtectionCrowdStrikeRuntime AI agent protection and shadow AI discovery
Microsoft Entra Agent IDMicrosoftAgent identity within Microsoft enterprise perimeter
Anthropic Trustworthy Agents / Zero Trust for AIAnthropic (Apr–May 2026)Model-layer safety and zero trust principles for agents
AIUC-1AI Underwriting Consortium / Lovable (May 2026)Six control families for coding agents
ASFJeff SutherlandAgent Security Framework
MATRAAcademicModel the attack surface of agentic AI systems
CoSAILinux FoundationCoalition for Secure AI - working groups and guidance
C2PAContent ProvenanceOrigin and edit history of media content
EU AI ActEuropean Commission (Aug 2026)Risk-based requirements for AI systems in the EU
ISO/IEC 42001ISOCertifiable AI management system; 12–18 month process
DORAEUDigital operational resilience for financial services
NIS2EUNetwork and information security across critical sectors

Every framework in this landscape tells an enterprise what their AI security posture should be.

Proof Benchmark tells them - and anyone who asks - what it actually is.

Your product should be able
to prove what it claims.

Submit a versioned build for evaluation against the Gauntlet. Pass or fail, results publish to ProofRegister - public, tamper-resistant, permanently on record. If scores meet PESA thresholds, a PROOFstamp is issued.

Request Evaluation What Is PROOFstamp?