v1.2-UNIFIED-ENTERPRISE REPORT

Aegis SDK Technical Audit & Benchmarks

Empirical security and latency evaluation of the Aegis Multi-Layer AI Security & Governance Engine across 51 total executions under OWASP 2025 standard test conditions.

Core Engine: Aegis Security SDK v0.5.5Python 3.13.15OWASP Top 10 LLM (2025)
Total Evaluated Executions
51
27 Telemetry Traces + 24 OWASP Suite Scenarios
Threat Defense Score
97.5%
39 / 40 Malicious Adversarial Prompts Blocked
Precision & Accuracy
90.9%
10 / 11 Legitimate Control Queries Allowed
False Positives / Negatives
1.96% / 2.5%
1 False Positive & 1 False Negative across 51 runs
Local Pre-Filter Overhead
< 1.0 ms
Sub-millisecond CPU heuristic regex interception
Mean Threat Interception
584.2 ms
Fastest interception: 513.4 ms (Base64 payload)

Dataset Breakdown

Dataset A (Telemetry Traces) vs Dataset B (Automated OWASP Suite)

51 Runs Unified
Evaluation MetricDataset A (Telemetry Traces)Dataset B (Automated Suite)Unified Benchmark Total
Total Evaluation Test Runs272451 Runs
Adversarial Threats Evaluated211940 Malicious Prompts
Attacks Intercepted & Blocked20 / 21 (95.2%)19 / 19 (100%)39 / 40 (97.5%)
Benign Control Prompts6511 Control Queries
Benign Prompts Allowed6 / 6 (100%)4 / 5 (80.0%)10 / 11 (90.9%)
Mean Interception Latency~612 ms~584 ms~598.2 ms

Threat Domain Coverage & Accuracy Matrix

Categorized according to OWASP Top 10 for LLM Applications (2025 Standard)

OWASP CodeCategory DescriptionPrompts TestedInterceptedBlock RatePrimary Threat Flags Triggered
LLM01Direct System Prompt Override88100%
prompt_injectionunauthorized_access
LLM02Sensitive Data Exfiltration66100%
sensitive_data_accessdata_breach
LLM01Jailbreak & Payload Generation6583.3%
malware_creationjailbreak
LLM06Privilege Escalation & Auth Bypass44100%
unauthorized_accessjailbreak
LLM08Destructive Operations (SQL/OS)44100%
destructive_actionprompt_injection
SEC-SOCSocial Engineering & Fraud44100%
financial_fraudphishing
SEC-ROLEPersona & Authority Hijacking44100%
unauthorized_accessmalicious_action
SEC-OBFEncoded & Obfuscated Payloads44100%
potential_malicious_codeprompt_injection
CTRL-SAFEBenign Control Suite111090.9% (Allowed)
Clean Execution - 1 Warning

Multi-Layer Execution Pipeline Efficiency

Two-stage execution pipeline designed to eliminate latency and API costs for enterprise workloads.

1

Stage 1: CPU Heuristic Pre-Filter

< 1.0 ms Overhead

Instant pattern matching for known dangerous keywords, system overrides, and regex guardrails. Drops malicious inputs locally without consuming cloud tokens or causing network latency.

2

Stage 2: Layer 1 & 2 Policy Analysis

584.2 ms Mean Latency

Contextual evaluation against custom natural language policies. Evaluates tool permissions, risk levels, and stateful human approval gates.

Threat Indicator Taxonomy

8 threat flag categories automatically tracked in telemetry logs

prompt_injectionLayer 1

Direct instruction override and persona manipulation

unauthorized_accessLayer 1/2

Unauthorized namespace or system role access requests

sensitive_data_accessLayer 2

Attempts to access environment variables, connection strings, or PII

destructive_actionLayer 2 (HITL)

SQL DROP/TRUNCATE/DELETE statements and destructive OS commands

malware_creationLayer 1

Prompts requesting ransomware, keyloggers, or exploit payloads

financial_fraudLayer 2

Impersonation scripts targeting financial authorization

phishingLayer 1

Social engineering SMS/email scam generation

potential_malicious_codeLayer 1

Encoded or obfuscated instructions (Base64/Hex)

Technical Due Diligence & Assessment

Enterprise Reliability

Across 51 total evaluation runs across two independent test phases, Aegis demonstrated an empirical 97.5% threat interception score (39/40 attacks blocked) and a 90.9% precision score on benign control inputs.

Cost-Effective Defense

By intercepting over 97% of threat patterns at Layer 1/2, Aegis prevents malicious traffic from incurring downstream LLM inference costs or exhausting API quota limits.

Zero-Trust Compliance

Local execution capabilities ensure that sensitive policy evaluations occur on-device or within private VPC boundaries, fulfilling strict financial (BFSI) and enterprise compliance standards.