Files
plugin-registry/plugins/guardrails
ixoblakp 8dec8f6038 Initial Arka plugin registry: Official plugins mirror + Arka-signed index
11 plugins from github.com/librefang/librefang-registry plugins/.
index.json / index.json.sig are signed with Arka's Ed25519 key
(not upstream stats.librefang.ai). Private key is not in this repo.
2026-09-01 10:22:07 +03:00
..

guardrails

Safety filter plugin that detects potentially harmful content patterns in user messages and injects warning memories into agent context. Uses only Python stdlib regex -- no external dependencies.

Detection Categories

Category Examples Memory Tag
PII Email addresses, phone numbers, SSNs, credit card numbers [guardrails:pii]
Prompt injection "ignore previous instructions", "you are now", "system prompt:" [guardrails:injection]
Credentials password=, api_key=, secret=, token=, PEM private keys [guardrails:credential]

Hooks

Hook Script Description
ingest hooks/ingest.py Scans user messages for harmful patterns and returns warning memories

How It Works

When a user message arrives, the ingest hook runs all pattern checks against it. For each detected issue a memory is returned with the category tag and a recommendation for the agent (e.g. "avoid echoing PII", "maintain original instructions"). If nothing is detected the plugin returns an empty memories list.

All patterns use word boundaries and anchoring to minimise false positives on casual conversation.

Usage

Installed automatically when enabled in agent configuration.