AIPWN / DECISION METHODS
Agent Security rule decisions: four reproducible examples
The examples below run the local deterministic layer with no semantic provider. They show actual rule behavior, not an accuracy or latency benchmark for the hosted API.
01 / METHOD
What was run
Four inputs pass through API schema validation → collectTextFields → runDeterministicRules → decide with an empty semantic-signal set. High-severity findings produce block, medium findings review, and the clean example allow. This is a behavior sample for the named rule and policy versions, not a statistical evaluation on an attack corpus.
The hosted endpoint also handles semantic signals, authentication, rate limits and failures, so a full API response can have additional risk codes or a different decision. Observe Beta decisions are advisory and do not stop a caller's tool execution.
02 / EXAMPLES
Edit an input and rerun the rules
This page runs API schema validation and deterministic rules in your browser. No sign-in or input upload is needed. The hosted endpoint also handles semantic signals and account state.
01 · Instruction override in retrieved text
Rule-only decision: block
Risk codes: INSTRUCTION_OVERRIDE, SYSTEM_PROMPT_EXTRACTION
02 · Tool request to a link-local address
Rule-only decision: block
Risk codes: SSRF_PRIVATE_TARGET
03 · Tool outside the capability map
Rule-only decision: review
Risk codes: UNKNOWN_TOOL_CAPABILITY
04 · Ordinary support text
Rule-only decision: allow
Risk codes: []
03 / LIMITS
Reproduce and apply
The four initial decisions are calculated by the current rule code in this page. Edit the JSON to see how risk codes change. The displayed identifiers track the published rule and policy versions.
These examples do not measure hosted end-to-end latency, false positives, false negatives or semantic-signal quality. Those metrics need a documented corpus, labeling method, sample count, test time and API configuration. Rules can miss paraphrased or novel attacks and can flag ordinary text.
Send the relevant event before a tool call, inspect the returned risk codes, and let your application decide when to request human review or stop execution.