Honest comparison

ContextWall vs everything else

Side-by-side against LlamaFirewall, NeMo Guardrails, Snyk agent-scan, PromptFoo, and Lakera. Where each wins, where each doesn't.

ContextWall Preflight

Safety preflight — release decision for the agent-as-a-system, plus runtime firewall.

visit

Unit of analysis

The running agent's behavior (model-boundary, trace-boundary, component-exposure).

Deployment

OSS CLI (Apache 2.0). Runtime daemon. Optional hosted scorecard (v1.5).

Where it wins

Only tool that returns a release gate + evidence-backed fix pages. Redacts secrets from evidence by default. CI-first exit-code contract.

Where it doesn't

Not a state-of-the-art classifier. Not an eval framework for training-time iteration. Won't tell you if your agent's answers are correct.

Meta LlamaFirewall / PromptGuard 2

Free classifier + guardrail primitives from Meta.

visit

Unit of analysis

Individual message text.

Deployment

Open-weights model + Apache-adjacent library.

Where it wins

SOTA classification accuracy in Meta's evals. Actively updated. Zero cost. Backed by Meta's distribution.

Where it doesn't

No release decision. No fix pages. No scorecard format. No runtime enforcement plane. History (React vs JS libs, PyTorch vs TensorFlow) says: competing head-to-head with a Meta free tool on classifier accuracy is the losing move.

Nvidia NeMo Guardrails

Programmable dialog rails via the Colang DSL.

visit

Unit of analysis

Conversation state machines.

Deployment

OSS library, Python-only.

Where it wins

Very expressive dialog control. Nvidia distribution. Composes well with a classifier underneath (e.g. Llama Guard).

Where it doesn't

Not a security scanner. Not a scorecard. Higher setup cost — you write Colang programs. Different job than Preflight.

Snyk agent-scan (ex-Invariant)

Static scanner for MCP servers, agents, skills, tools, prompts, resources.

visit

Unit of analysis

MCP component inventory (tools + resources on a server).

Deployment

OSS CLI (Apache 2.0), 2.9k stars, backed by Snyk's 4.5M-developer distribution.

Where it wins

Owns the MCP-component-scanning wedge. Snyk brand. Weekly releases. Explicit tool-poisoning / tool-shadowing / rug-pull detection.

Where it doesn't

Scans components, not agent behavior. Doesn't grade the agent-as-a-system. Doesn't produce a release decision on 'is my agent safe to ship'. Doesn't enforce at runtime.

PromptFoo

Eval framework for LLMs and prompts.

visit

Unit of analysis

Individual prompt / response pairs, run through user-defined assertions.

Deployment

OSS CLI, ~5k stars.

Where it wins

Flexible eval assertions. Great for prompt iteration and A/B testing prompts. Big install base.

Where it doesn't

Not a release gate. Not opinionated about which attacks matter. No fix pages. Different mental model — for devs tuning prompts, not devs asking 'is this safe to ship?'

Lakera Guard (Check Point)

SaaS API for prompt-injection / PII / data-leakage detection.

visit

Unit of analysis

Individual request/response text.

Deployment

Proprietary SaaS API. Enterprise sales motion. Acquired by Check Point Sept 2025.

Where it wins

Mature detection. Enterprise support. Now part of Check Point's larger security portfolio.

Where it doesn't

API-first, not agent-as-system. No OSS core. Sends your traffic to a third party. Very different pricing and deployment posture than a CLI.

Notes on honesty

  • We do not claim ContextWall's detector beats Meta PromptGuard 2 on classifier accuracy. History says that's the losing bet. The runtime firewall's detector is 'good enough' — the differentiation is what wraps it: the opinionated grader, gates, boundaries, redaction, and shareable scorecard format.
  • If your job is scanning MCP servers for vulnerable tools, Snyk agent-scan is the tool. Preflight's --mcp adapter is deliberately conservative (never invokes tools) — good for a safety-first read, not a replacement for Snyk's inventory scanner.
  • If your job is iterating on prompts and running A/B evals, PromptFoo is the tool. Preflight is for a different question: 'is this ready to ship?'
  • The categories overlap. Nothing wrong with running Preflight + Snyk + a classifier + evals — they answer different questions.

Run it against your agent

One command, no API keys, no config.

$ uvx contextwall check --openai http://localhost:11434/v1