Skip to main content
Declaw provides a layered security model. Every sandbox has a SecurityPolicy that defines exactly what kinds of outbound traffic are allowed, what PII gets redacted, and which requests get audited. All enforcement happens transparently in the security proxy running inside the sandbox — your agent code requires no modifications.

SecurityPolicy structure

Enforcement pipeline

All outbound traffic from the sandbox passes through a 6-stage pipeline before reaching the internet. On the response path, the body is not blocked or injection-scanned — it’s passed through to the agent. What runs on responses: inbound transformation rules, then PII rehydration (restoring original values from the session redaction map), plus capture of untrusted content as session context (used by the optional LLM judge for session-aware indirect-injection detection), and audit logging.

Stage descriptions

IP and CIDR rules are enforced at the kernel level via iptables. This is the fastest path — no userspace proxy overhead for purely IP-based rules. deny_out entries become DROP rules; allow_out IP/CIDR entries become ACCEPT rules with higher priority.
When domain names appear in allow_out or deny_out, all TCP traffic is redirected through the per-namespace TCP proxy. The proxy inspects the TLS SNI field (port 443) or HTTP Host header (port 80) to determine the destination domain. Wildcard patterns like *.openai.com are supported.
When PII scanning or transformation rules are enabled, the proxy performs TLS interception at the edge proxy. A per-sandbox CA certificate is generated at sandbox creation and injected into the VM trust store. The proxy terminates TLS, inspects the plaintext body, and re-encrypts to the real destination. This stage is skipped entirely when no body inspection is needed.
Scans outbound request bodies for PII and prompt injection; on responses it rehydrates PII (and captures content for indirect-injection provenance) but does not block them.PII scanning: Regex patterns cover structured PII (SSN, credit card with Luhn validation, email, phone). When the optional Guardrails Service is deployed, it adds ML-based NER for unstructured PII (person, location, passport, driver’s license). Three actions are available: redact (replace with a token), block (reject the request), or log_only (pass through and audit). The redaction map is stored per-session so response bodies can be rehydrated.Injection defense: Configurable sensitivity threshold. When the Guardrails Service is deployed, an ML classifier plus an LLM judge scores the content. Actions: block (reject) or log_only (pass through and audit). Injection scanning is opt-in per domain — it runs only on the hosts listed in domains (an empty list means no injection scanning), so scope it to your model endpoint. See Prompt Injection Defense.On the response path, the Transform Engine applies inbound rules and PII rehydration restores original values; the response is not injection-blocked.
Applies TransformationRule regex patterns to request or response bodies. Rules are direction-aware: outbound rules apply to requests, inbound rules apply to responses. Useful for stripping API keys from outbound headers or removing injection patterns from inbound content.
Records lifecycle events (vm_created, vm_killed, …) and, when audit is enabled, network decisions (egress_allowed, egress_blocked) to the platform audit log. Request/response bodies are not recorded. Retention is 7 days platform-wide; opt out per sandbox with AuditConfig(enabled=False).

Composability

Each security component is independent. You can enable any combination:
TLS interception (Stage 3) activates automatically when pii.enabled=True or transformations are configured. It remains off when only network policies or audit logging are used, so there is no TLS overhead for pure network restriction use cases.

Shorthand forms

Several fields accept shorthand values for common configurations:

Security sections