Injection firewall
A regex fast path sheds high-volume attacks at zero ML cost; an embedding-based semantic classifier judges the subtle ones. Three operating modes tune recall against false positives.
An AI firewall stops the trick before your agent reads it. A PolicyGuard contract on Monad stops the money if anything gets through. Mera passkeys give each agent its identity and sealed memory.
99.0% AdvBench strict|41.88 ms p95 under attack|207/207 Hardhat tests|1,490+ Python suites
Architecture note: This sandbox runs the fast-path screen and de-obfuscation in your browser. The semantic classifier, output redaction and the on-chain PolicyGuard check run in the full gateway; try them in the Attack Lab.
AdvBench block rate · strict
benchmark run · aug 2026
p95 latency while blocking
chaos report · aug 2026
real-time security controls
in one gateway pass
OWASP LLM Top 10 (2025)
risks covered · test snapshot
DISCLOSURE: Compliance and audit artifacts are internally verified empirical test suites (3,211 prompts across 8 corpora, 1,490+ Python & 207 Hardhat tests, 184 Mera Passkey PRF assertions), not third-party statutory certifications.
Each layer is independently tested and artifact-documented. Together they close the attack classes that ship in the wild today.
A regex fast path sheds high-volume attacks at zero ML cost; an embedding-based semantic classifier judges the subtle ones. Three operating modes tune recall against false positives.
Morse, Braille steganography, Base64, hex, binary, ROT13, homoglyphs — normalized before classification so payloads can't hide behind an encoding layer.
Responses scanned before egress: PII redacted, XSS/SQL/shell fragments blocked, system-prompt leakage caught, optional watermarking applied downstream.
Rate limiting that closes instead of opens under failure, chaos-tested degradation, and SLO checks that run in CI — when parts die, the gateway fails safe.
The execution layer decides in milliseconds whether traffic is safe. The trust layer makes an agent's identity and track record independently verifiable.
Every figure on this site cites its dated repository artifact — and the proof page publishes full benchmark tables, not just highlights.
3,211 prompts across 8 public datasets, August 2026. AdvBench 99.0% strict, JBB PAIR+GCG 98.7% — and Do-Not-Answer's 57.1% published too, because selective reporting is how trust dies.
One external review of the off-chain gateway (March 2026, summary published, zero open critical/high). Internal off-chain, contract and architecture audits July–September 2026, including the Mera passkey enclave audit (166/166 assertions). The Monad contracts have not been externally audited yet.
Nine contracts live on Monad Testnet, each checked live in the Explorer proof ledger. State anchoring runs simulated by default. No mainnet claims.
The Explorer at aiguardian.dev/explorer runs real attacks against GuardianAI's gates and calls the deployed contracts directly from your browser.
Pick an attack (hidden instructions, a scam invoice, a forged approval, an impostor agent) and watch both gates refuse it: the off-chain firewall first, then PolicyGuard on Monad, with the real call and revert shown.
Launch an attack ↗Every contract's code and live settings, read from Monad testnet each time the page loads, plus a real approved payment.
Open the proof ledger ↗Connect a wallet (read-only), see agent ID cards, check any address against the on-chain scam list, and read the payment guard's live settings.
Open the Console ↗