Private and on-premises AI · 2 min read

Auditing a local AI agent before production release

Agent acceptance must test the complete chain from evidence retrieval and model decision to authorised action, actual outcome and reconstructable audit record.

Freeze the candidate build

Inventory model origin, runtime, configuration, instructions, retrieval, sources, tools and external dependencies. Automatic changes make test evidence ambiguous.

Freeze every component and configuration used in evaluation so the released system can be tied to the recorded evidence.

Test from threats

Cover cross-user leakage, malicious instructions in content, tool-parameter abuse, replay after failure, fabricated evidence, unavailable services and resource exhaustion.

Derive cases from data disclosure, hostile content, tool misuse, replay, fabricated sources, dependency failure and resource exhaustion risks.

Separate measurements

Score retrieval, factual support, refusal, tool choice, arguments, operation outcome and latency independently, with strict thresholds for severe failures.

Report retrieval, grounding, refusal, tool selection, arguments, outcome and latency independently, with separate limits for severe failures.

Accept operations too

Verify monitoring, restore, credential rotation, user revocation, source deletion, tool shutdown, model rollback and incident procedure, then rerun affected tests after changes.

Verify monitoring, recovery, credential rotation, access revocation and rollback, then rerun affected cases after every material change.

Evidence and sources

Continue with the underlying material

pommeDeTerre

Need an estimate for your project?

Tell us about the project. We will break it into stages and explain the budget drivers.

View pricing