White Paper Version 2
WhiteHaX AI-ASM provides an independent, repeatable way to validate how deployed AI systems behave under adversarial, malformed, policy-violating, high-volume, and tool-abuse conditions. Version 2 documents the current endpoint and profile model, expanded AI-specific DoS testing, protocol-aware MCP assurance, custom prompt and document generation, evidence processing, remediation, DevSecOps outputs, and service-provider operating patterns.
|
Deployment |
Execution |
Evidence |
|---|---|---|
|
Cloud-managed with on-premise application support |
Console, REST API, CLI, schedules, and CI/CD patterns |
Readiness, failures, response metrics, recommendations, PDF or HTML, JSON, and SARIF |
|
Section |
Purpose |
|---|---|
|
1 Executive Summary |
Business and technical position |
|
2 Platform Scope |
Targets, endpoint definitions, and profiles |
|
3 Architecture and Data Flow |
Inputs, testers, analysis, and outputs |
|
4 Security and Safety Test Domains |
Prompt, document, API, model, and compliance coverage |
|
5 Regulation Compliance Validation |
Seven framework packs and evidence-report schema |
|
6 AI Resilience Performance and Cost |
DoS, response-time, and cost-risk validation |
|
7 MCP Assurance |
Protocol-aware safe and intrusive testing |
|
8 Custom Test Generation |
Customer-specific prompts, documents, and policies |
|
9 Evidence Scoring and Remediation |
Result normalization, readiness, and recommendations |
|
10 Automation and DevSecOps |
REST, CLI, schedules, CI/CD, JSON, and SARIF |
|
11 Deployment and Operating Controls |
Safe execution, separation, and retesting |
|
12 Enterprise and Service Provider Use |
Teams, MSSPs, and pen testers |
|
13 Sources and Claim Boundaries |
Website, standards, help, and code evidence |
1 Executive Summary
Generative AI applications expose a combined attack surface: model behavior, prompts, retrieval context, uploaded files, APIs, agent tools, MCP services, identity and authorization boundaries, performance, and downstream cost. Traditional application testing remains necessary, but it does not by itself measure whether an AI deployment follows policy, refuses unsafe requests, contains sensitive data, survives AI-specific abuse, or keeps tool use within approved boundaries.
WhiteHaX AI-ASM tests the deployed system through its exposed interfaces. Operators define endpoints, select repeatable profiles, use built-in or custom test inputs, execute through interactive or automated paths, and receive evidence suitable for engineering, risk, compliance, and retesting. The platform separates testing from runtime-control sales, which supports independent validation across model providers, agent frameworks, gateways, WAFs, and guardrails.
The current product position has two complementary outcomes. SecureAI testing concentrates on security, safety, compliance, data exposure, and abuse. OptimalAI testing measures responsiveness, resilience, resource pressure, tail latency, and cost-amplification conditions. Together they help teams evaluate whether an AI service is secure, policy-conformant, responsive, resilient, and cost-controlled.
An endpoint records how WhiteHaX connects to the target. A profile records what to test, which inputs and parameters to use, and how the run should be controlled. This separation lets teams reuse a validated connection definition across smoke, regression, compliance, resilience, and retest profiles without mixing their risk boundaries.
|
Endpoint family |
Purpose |
|---|---|
|
AI application or agentic API |
Tests a deployed application workflow through scripts or JSON request and response mappings. |
|
Model endpoint |
Tests a first-party or third-party model interface or gateway directly. |
|
MCP server endpoint |
Tests protocol behavior, tools, resources, prompts, transport, and abuse handling. |
|
Regulation-compliance endpoint |
Runs policy or regulatory cases through a compliance-specific mapping. |
|
Profile family |
Primary objective |
|---|---|
|
Prompts and Docs |
Prompt safety, jailbreaks, leakage, unsafe output, and malicious-document or RAG behavior. |
|
AI-specific DoS |
Availability, resource pressure, cost amplification, abusive concurrency, and malformed workload behavior. |
|
AI Model Attack Surface |
Model-facing behavioral and security risks independent of a larger application workflow. |
|
MCP Attack Surface |
Protocol-aware safe checks and separately authorized intrusive MCP tests. |
|
Response Time Optimization |
Response-time measurement, performance baselining, and regression analysis. |
|
Regulation Compliance |
Selected customer or built-in policy and regulatory datasets. |
The control plane accepts operator and automation requests. Endpoint configuration, profile configuration, test files, and adapters are resolved into a run context. Specialized runners execute prompt and document, DoS, RTO, and MCP assessments. Response analysis converts raw target behavior into structured evidence. Readiness and recommendation components then package the results for human and machine consumers.
Figure 1 WhiteHaX AI-ASM platform architecture
|
Layer |
Components |
Responsibility |
|---|---|---|
|
Access |
Console, REST API, CLI, schedule, CI/CD |
Starts, monitors, and retrieves authorized assessments. |
|
Configuration |
Endpoints, profiles, scripts, JSON mappings, test parameters |
Defines the target, test family, execution boundaries, and response extraction. |
|
Inputs |
Built-in libraries, uploaded prompts and documents, generated files, MCP JSON, DoS and RTO cases |
Provides reproducible test cases and customer-specific context. |
|
Execution |
Prompt and document runner, model tester, DoS runner, RTO runner, MCP runners |
Sends cases, manages concurrency, and captures target behavior. |
|
Analysis |
Response analyzer, pass or fail logic, readiness calculation, metrics |
Classifies evidence and calculates overall and subcategory posture. |
|
Recommendation |
Prompt and document, DoS, MCP, and response-time generators |
Maps failure type and severity to remediation actions and retest scope. |
|
Output |
Dashboard, PDF or HTML, JSON, SARIF |
Delivers evidence to security, engineering, audit, and DevSecOps systems. |
Prompt testing covers direct injection, indirect injection, jailbreaks, hidden intent, backdoor triggers, model hijacking, contextual drift, multi-step malicious intent, bias and influence attempts, and unsafe changes in model behavior. Expected-result patterns and benign controls help distinguish a true bypass from a safe refusal or an unrelated response.
Leakage tests attempt to expose PII, protected or confidential information, proprietary context, system instructions, and other restricted data. Output analysis can also test toxic content, biased recommendations, and unsafe disclosures. Customer-specific signatures should be used when generic prohibited terms cannot establish whether a business workflow failed.
Document-aware testing exercises uploaded Office files, PDFs, images, QR codes, embedded objects, links, metadata, malformed structures, hidden instructions, and retrieval-grounded content. These cases test the complete ingestion and retrieval path, not only the model prompt field. Generated malicious documents can extend the built-in library with authorized customer scenarios.
API tests cover malformed and invalid requests, token or key abuse, replay and permission-bypass attempts, anomalous access patterns, access-right violations, scraping, and exfiltration behavior. Application authentication and tenant-routing controls remain part of the target architecture; WhiteHaX supplies repeatable probes and evidence rather than replacing those controls.
The platform supports privacy, regulated-data, AI-governance, assurance, and application-security cases through versioned test packs and customer-defined mappings. WhiteHaX maps technical observations to selected requirements and retains the underlying evidence; the dedicated section below defines the seven named frameworks and the evidence report.
Figure 2 Validation remediation and retest flow
WhiteHaX regulation-compliance profiles turn approved prompts, documents, API calls, model interactions, RAG content, agent actions, and MCP cases into repeatable control tests. A framework pack identifies the reference version, applicability assumptions, validation objective, test cases, evidence rules, and report mappings. The organization selects the obligations and controls in scope; WhiteHaX validates observable technical behavior and records gaps for remediation and retesting.
|
Framework |
Technical validation focus |
Evidence report examples |
|---|---|---|
|
GDPR |
Privacy and personal-data handling: minimization, purpose and access boundaries, PII disclosure, retention or deletion behavior, and data-subject workflow tests. |
Mapped test IDs, request and response evidence, exposed data classes, expected control behavior, gap and retest result. |
|
EU AI Act |
Role- and risk-class-aware tests supporting risk management, data governance, technical documentation, logging, transparency, human oversight, accuracy, robustness, and cybersecurity evidence. |
Applicable article or obligation mapping, system role and risk assumption, test result, limitation, owner, and remediation trail. |
|
OWASP GenAI LLM Top 10 |
Versioned coverage for the current OWASP GenAI LLM Top 10, including prompt injection, sensitive-data disclosure, unsafe output handling, excessive agency, supply-chain or poisoning risks, and resource abuse. |
OWASP risk ID and release, attack case, observed behavior, severity, mitigation evidence, and regression status. |
|
NIST AI RMF |
Evidence organized across Govern, Map, Measure, and Manage, with tests for validity, safety, security, resilience, accountability, transparency, privacy, and harmful-bias risk. |
AI RMF function or outcome mapping, measurement evidence, risk decision, treatment, owner, and monitoring or retest record. |
|
SOC 2 |
Technical readiness evidence aligned to the applicable Trust Services Criteria for security, availability, processing integrity, confidentiality, and privacy. |
Control objective mapping, test procedure, result, exception, evidence reference, remediation owner, and retest; not a CPA attestation. |
|
HIPAA |
PHI and ePHI leakage, access boundaries, minimum-necessary behavior, integrity, availability, auditability, and administrative or technical safeguard validation relevant to the target. |
Safeguard or requirement mapping, PHI-safe test evidence, access and disclosure result, exception, mitigation, and validation date. |
|
PCI DSS 4.0.1 |
Cardholder-data exposure, prompt, output, document, API, identity, logging, and security-control behavior within the assessed AI service and connected workflow. |
Requirement mapping, scoped component, test evidence, status, compensating-control note, remediation, and retest; not a QSA assessment. |
|
Report area |
Required content |
|---|---|
|
Scope and mapping |
Framework and version, requirement or control reference, applicability, system role or risk class, target endpoint, environment, model, and profile version. |
|
Test specification |
Stable test ID, validation objective, input or artifact reference, expected behavior, evidence rule, execution parameters, and authorization boundary. |
|
Observed evidence |
Timestamp, request and normalized response, artifact hash or reference, actual behavior, pass, fail, not tested, or not applicable status, severity, and rationale. |
|
Closure trail |
Finding and limitation, recommended remediation, owner, exception or compensating control, retest date and result, residual risk, and reviewer sign-off field. |