Shadow AI · Insecure Defaults · Indirect Injection

2026 Enterprise AI Agent Security Scorecard: Insecure Defaults & Indirect Prompt Injection Risks

Independent evaluation matrix for the top 10 agent runtimes — measuring out-of-the-box defense against indirect hijacking, tool abuse, injection in tool parameters, and unsafe multi-agent delegation. Nexus Shield runtime control restores 99.4%+ mitigation at <12ms intercept latency.

80% of enterprises deploy AI agents with out-of-the-box defaults. 86% of those agents are vulnerable to indirect prompt hijacking.

Average framework defense without runtime control: 19.8% (931% range across vectors).

Interactive evaluation matrix

Expand any framework for per-vector comparison: insecure default (FAIL/PARTIAL) vs protected with Nexus Shield (PASS · ~12ms), intent divergence, READ_ONLY revocation, and MCP-SEC-SCORE evidence IDs.

Showing 10 frameworks · lens: All attack vectors

FrameworkOut-of-the-boxWith Nexus ShieldDivergence (OOTB)CapabilityEvidence
CrewAI
16.5% defendedFail
99.5% mitigatedPass · 9.8ms
0.71READ_ONLYMCP-SEC-SCORE-crewai-aggregate

Vector

Indirect Prompt Injection via Untrusted Input

Insecure default

14% defended

Fail

Nexus Shield

99.4% mitigated

Pass · 9.2ms

Divergence · Revocation · Evidence

Δ 0.710.03

READ_ONLY

MCP-SEC-SCORE-crewai-indirect_injection

Vector

Tool Abuse & Excessive Agency

Insecure default

18% defended

Fail

Nexus Shield

99.6% mitigated

Pass · 10.1ms

Divergence · Revocation · Evidence

Δ 0.690.03

READ_ONLY

MCP-SEC-SCORE-crewai-tool_abuse

Vector

Unsanitized Tool Arguments

Insecure default

22% defended

Partial

Nexus Shield

99.8% mitigated

Pass · 8.4ms

Divergence · Revocation · Evidence

Δ 0.670.03

READ_ONLY

MCP-SEC-SCORE-crewai-unsanitized_tool_args

Vector

Unsafe Inter-Agent Delegation

Insecure default

12% defended

Fail

Nexus Shield

99.1% mitigated

Pass · 11.6ms

Divergence · Revocation · Evidence

Δ 0.720.03

READ_ONLY

MCP-SEC-SCORE-crewai-inter_agent_delegation

LangChain / LangGraph
19.8% defendedFail
99.3% mitigatedPass · 9.9ms
0.70READ_ONLYMCP-SEC-SCORE-langchain-langgraph-aggregate
AutoGen
13.3% defendedFail
99.5% mitigatedPass · 9.7ms
0.74READ_ONLYMCP-SEC-SCORE-autogen-aggregate
LlamaIndex
22.3% defendedPartial
99.4% mitigatedPass · 10.0ms
0.69READ_ONLYMCP-SEC-SCORE-llamaindex-aggregate
OpenAI Assistants
28.3% defendedPartial
99.2% mitigatedPass · 9.6ms
0.66READ_ONLYMCP-SEC-SCORE-openai-assistants-aggregate
Semantic Kernel
17.5% defendedFail
99.5% mitigatedPass · 9.8ms
0.71READ_ONLYMCP-SEC-SCORE-semantic-kernel-aggregate
Haystack
24.5% defendedPartial
99.3% mitigatedPass · 9.9ms
0.68READ_ONLYMCP-SEC-SCORE-haystack-aggregate
DSPy
10.5% defendedFail
99.6% mitigatedPass · 9.5ms
0.75READ_ONLYMCP-SEC-SCORE-dspy-aggregate
SuperAGI
13.5% defendedFail
99.5% mitigatedPass · 9.7ms
0.73READ_ONLYMCP-SEC-SCORE-superagi-aggregate
MCP Native SDKs
26.5% defendedPartial
99.4% mitigatedPass · 10.1ms
0.67READ_ONLYMCP-SEC-SCORE-mcp-native-sdks-aggregate
OOTB defense (avg)

19.8%

9–31% range

Nexus Shield mitigation

99.4%

<12ms intercept p50

Evidence standard

MCP-SEC-SCORE

Per-vector cryptographic bundles

CISO outreach & local audit

Share the executive summary with security leadership or reproduce the full scorecard in your CI pipeline — no Nexus Shield account required.

Run local audit (Docker)

docker run --rm ghcr.io/baturhantasdelen-sudo/harness:latest --eval-scorecard