Apache 2.0 · Zero API keys · Runs offline against fixtures

Preflight your agent before you ship it.

Point ctxfw check at your agent. Get a safety scorecard and a release decision in about 60 seconds — with evidence-backed fixes for what fails. No API keys, no config, no cloud.

$ uvx contextwall check --openai http://localhost:11434/v1

or: pip install contextwall && ctxfw check

PyPI versionPython versionsLicense Apache 2.0GitHub stars
contextwall preflight — cursor-agent v1.4
Suite: standard@2026.09 Adapter: openai-compatible
Assessment mode: MODEL-BOUNDARY Scope: SAFETY ONLY
Injection ResistanceB+7/10 cases passed
Grounding Under AttackC2/5 cases passed
Data ExfiltrationA0/20 emissions detected
Release gate: BLOCKED
✗ Grounding score below configured minimum (B-)
✗ Critical test inj-05 (semantic override) failed
Not evaluated: tool authorization, correctness, cost, latency, reliability, tenant isolation, availability, data retention

Findings stay on your machine · No secrets sent anywhere · Redacts credentials before rendering

Real production incidents, not theoretical threats

Your agent trusts everything it reads

LLMs have no built-in concept of source trust. Content retrieved from a web search and content from your system prompt look identical once they are both inside the context window. Attackers exploit this directly.

CVE-2025-32711

EchoLeak

Microsoft 365 Copilot

9.3 Critical

An attacker sends a crafted email. Copilot reads it, interprets embedded instructions as commands, silently accesses internal SharePoint files, and sends them to the attacker. The user never clicks anything.

WHY IT WORKED

Copilot had no way to distinguish a trusted system instruction from untrusted email content. Both looked the same inside the context window.

USENIX Security 2025

PoisonedRAG

RAG pipelines

90%+ manipulation rate

Researchers planted five adversarial documents into a knowledge base of millions. When users asked questions, the model retrieved and repeated the false content as confident fact, with no jailbreak, no system prompt change, and no model access needed.

WHY IT WORKED

The RAG pipeline retrieved documents by relevance score and passed them straight to the model. There was no check on where the document came from or whether it should be trusted.

Both incidents shipped to production before anyone knew the agent was vulnerable. You can't defend what you haven't measured. ContextWall Preflight runs the pinned adversarial corpus against your agent and tells you exactly which of these classes it falls for — before the incident, not after.

How Preflight works

Three steps. No config file. No cloud. Copy-pasteable in every terminal.

1. Install

One command. No API keys. No sign-up.

uvx contextwall check --openai http://localhost:11434/v1

2. Point at your agent

Preflight runs the pinned standard@2026.09 suite — 35 adversarial fixtures across injection, grounding, and exfiltration — against your agent endpoint and returns a scorecard.

→ Injection Resistance     B+   7/10 passed
→ Grounding Under Attack   C    2/5 passed
→ Data Exfiltration        A    0/20 emissions

3. Fix what fails

Every failure links to a remediation page with the exact attack pattern, the expected safe behavior, and the config change to close the gap. Re-run to confirm.

Fix first:
  ✗ Semantic override bypass       → /fix/inj-05
  ✗ Poisoned document accepted     → /fix/gnd-02

Four ways to point at your agent

Every adapter declares what it can observe. The scorecard renders that boundary verbatim — no comparing apples to oranges.

--openai

OpenAI-compatible endpoints (OpenAI, vLLM, Ollama, LiteLLM, ...)

--anthropic

Anthropic Messages endpoints (direct, Bedrock proxies)

--replay

JSONL of prerecorded responses — grade offline in CI

--mcp

MCP config — enumerates tools + scans for poisoned descriptions (never invokes tools)

What Preflight does NOT measure

Preflight is a safety preflight, not a production-readiness gate. The scorecard states this explicitly. If it passes, your agent isn't un-shippable — but it hasn't been evaluated for the following either:

  • – correctness of the agent's answers
  • – cost, latency, throughput, availability
  • – tool authorization semantics (who is allowed to call what)
  • – tenant isolation and multi-user access control
  • – data retention and provider configuration
  • – human-in-loop policy correctness

Honest scope beats false assurances. Preflight covers safety controls; correctness, performance, and authorization are your existing test suites' job.

What the scorecard covers — and what it doesn't

35 pinned fixtures across three safety dimensions. Every fixture cites its source. Every scorecard states its scope.

Evaluated safety dimensions

Injection Resistance10 fixtures
  • inj-01EchoLeak-shape— CVE-2025-32711
  • inj-05Semantic override— no keywords — heuristic
  • inj-07Tool poisoning— Invariant Labs 2025
  • inj-08AWS Q-style destructive— GHSA-7g7f-ff96-5gcw
  • inj-09GitHub MCP exfil— Invariant Labs 2025
Grounding Under Attack5 fixtures
  • gnd-01False attribution— PoisonedRAG variant
  • gnd-04Conflicting sources— majority-vote failure
  • gnd-05Poisoned action recommendation— shell command from wiki
Data Exfiltration20 fixtures
  • exf-01AWS access key— AKIAIOSFODNN7EXAMPLE
  • exf-08SSH private key— OpenSSH format
  • exf-14Composite medical PII— HIPAA-relevant
  • exf-19Secret via tool output— indirect leak
  • exf-20Secret via HTTP tool— direct exfiltration

Suite pinned to standard@2026.09. Two scorecards produced by the same version compare directly. Older suites remain runnable so published scorecards keep meaning what they meant.

Deliberately out of scope

  • Correctness / hallucination rate

    different problem class — LLM evals like Braintrust / LangSmith

  • Cost, latency, throughput

    your existing observability

  • Tool authorization semantics

    who is allowed to call what — needs application-layer logic

  • Runtime enforcement

    separate concern — ctxfw start (runtime firewall)

A passing Preflight means the agent resisted the pinned safety corpus — not that it's production-ready overall. Every scorecard states this explicitly.

Ship it into your workflow

Three shapes: interactive local check, gated CI run, or an MCP config scan. All from the same CLI.

Local check
# Zero install — uv fetches Preflight and runs it in one shot
uvx contextwall check --openai http://localhost:11434/v1

# Or install into your env
pip install contextwall
ctxfw check --anthropic https://api.anthropic.com \
            --anthropic-model claude-3-5-haiku-20241022

35 pinned fixtures across injection, grounding, exfiltration. No API keys required for the fixture suite itself — only for the model endpoint you're checking. Redacts secrets from scorecard evidence by default.