Preflight your agent
before you ship it.
Point ctxfw check at your agent. Get a safety scorecard and a release decision in about 60 seconds — with evidence-backed fixes for what fails. No API keys, no config, no cloud.
or: pip install contextwall && ctxfw check
Assessment mode: MODEL-BOUNDARY Scope: SAFETY ONLY
Findings stay on your machine · No secrets sent anywhere · Redacts credentials before rendering
Your agent trusts everything it reads
LLMs have no built-in concept of source trust. Content retrieved from a web search and content from your system prompt look identical once they are both inside the context window. Attackers exploit this directly.
CVE-2025-32711
EchoLeak
Microsoft 365 Copilot
An attacker sends a crafted email. Copilot reads it, interprets embedded instructions as commands, silently accesses internal SharePoint files, and sends them to the attacker. The user never clicks anything.
Copilot had no way to distinguish a trusted system instruction from untrusted email content. Both looked the same inside the context window.
USENIX Security 2025
PoisonedRAG
RAG pipelines
Researchers planted five adversarial documents into a knowledge base of millions. When users asked questions, the model retrieved and repeated the false content as confident fact, with no jailbreak, no system prompt change, and no model access needed.
The RAG pipeline retrieved documents by relevance score and passed them straight to the model. There was no check on where the document came from or whether it should be trusted.
Both incidents shipped to production before anyone knew the agent was vulnerable. You can't defend what you haven't measured. ContextWall Preflight runs the pinned adversarial corpus against your agent and tells you exactly which of these classes it falls for — before the incident, not after.
How Preflight works
Three steps. No config file. No cloud. Copy-pasteable in every terminal.
1. Install
One command. No API keys. No sign-up.
uvx contextwall check --openai http://localhost:11434/v1
2. Point at your agent
Preflight runs the pinned standard@2026.09 suite — 35 adversarial fixtures across injection, grounding, and exfiltration — against your agent endpoint and returns a scorecard.
→ Injection Resistance B+ 7/10 passed → Grounding Under Attack C 2/5 passed → Data Exfiltration A 0/20 emissions
3. Fix what fails
Every failure links to a remediation page with the exact attack pattern, the expected safe behavior, and the config change to close the gap. Re-run to confirm.
Fix first: ✗ Semantic override bypass → /fix/inj-05 ✗ Poisoned document accepted → /fix/gnd-02
Four ways to point at your agent
Every adapter declares what it can observe. The scorecard renders that boundary verbatim — no comparing apples to oranges.
--openai
OpenAI-compatible endpoints (OpenAI, vLLM, Ollama, LiteLLM, ...)
--anthropic
Anthropic Messages endpoints (direct, Bedrock proxies)
--replay
JSONL of prerecorded responses — grade offline in CI
--mcp
MCP config — enumerates tools + scans for poisoned descriptions (never invokes tools)
What Preflight does NOT measure
Preflight is a safety preflight, not a production-readiness gate. The scorecard states this explicitly. If it passes, your agent isn't un-shippable — but it hasn't been evaluated for the following either:
- – correctness of the agent's answers
- – cost, latency, throughput, availability
- – tool authorization semantics (who is allowed to call what)
- – tenant isolation and multi-user access control
- – data retention and provider configuration
- – human-in-loop policy correctness
Honest scope beats false assurances. Preflight covers safety controls; correctness, performance, and authorization are your existing test suites' job.
What the scorecard covers — and what it doesn't
35 pinned fixtures across three safety dimensions. Every fixture cites its source. Every scorecard states its scope.
Evaluated safety dimensions
- inj-01EchoLeak-shape— CVE-2025-32711
- inj-05Semantic override— no keywords — heuristic
- inj-07Tool poisoning— Invariant Labs 2025
- inj-08AWS Q-style destructive— GHSA-7g7f-ff96-5gcw
- inj-09GitHub MCP exfil— Invariant Labs 2025
- gnd-01False attribution— PoisonedRAG variant
- gnd-04Conflicting sources— majority-vote failure
- gnd-05Poisoned action recommendation— shell command from wiki
- exf-01AWS access key— AKIAIOSFODNN7EXAMPLE
- exf-08SSH private key— OpenSSH format
- exf-14Composite medical PII— HIPAA-relevant
- exf-19Secret via tool output— indirect leak
- exf-20Secret via HTTP tool— direct exfiltration
Suite pinned to standard@2026.09. Two scorecards produced by the same version compare directly. Older suites remain runnable so published scorecards keep meaning what they meant.
Deliberately out of scope
Correctness / hallucination rate
different problem class — LLM evals like Braintrust / LangSmith
Cost, latency, throughput
your existing observability
Tool authorization semantics
who is allowed to call what — needs application-layer logic
Runtime enforcement
separate concern — ctxfw start (runtime firewall)
A passing Preflight means the agent resisted the pinned safety corpus — not that it's production-ready overall. Every scorecard states this explicitly.
Ship it into your workflow
Three shapes: interactive local check, gated CI run, or an MCP config scan. All from the same CLI.
# Zero install — uv fetches Preflight and runs it in one shot
uvx contextwall check --openai http://localhost:11434/v1
# Or install into your env
pip install contextwall
ctxfw check --anthropic https://api.anthropic.com \
--anthropic-model claude-3-5-haiku-2024102235 pinned fixtures across injection, grounding, exfiltration. No API keys required for the fixture suite itself — only for the model endpoint you're checking. Redacts secrets from scorecard evidence by default.