Category archive

Evaluation and security are one discipline aimed at two failures.

Both questions are the same question asked twice: what does this system do when its input is wrong, and what does it do when its input is hostile. An evaluation harness that cannot block a release and a guardrail that only blocks keywords fail for the same reason, so they belong in one archive.

7 articlesEvaluation and Security

Articles in Evaluation and Security

Synthetic Evaluation Data That Finds Real Failures

How to generate, filter, diversify, and maintain synthetic cases without turning an evaluation suite into a mirror of its generator.

Privacy-Preserving AI Is a Dataflow Architecture

Purpose limitation, minimization, isolation, retention, redaction, and verifiable deletion across retrieval, models, tools, traces, and feedback loops.

Build an LLM Evaluation Harness That Can Block a Release

From task contracts and test slices to calibrated judges, regression budgets, and production feedback loops.

Hallucination Controls Belong in the System, Not One Prompt

A layered design for constraining claims, grounding evidence, verifying outputs, and abstaining when the system does not know.

Guardrails as Policy Enforcement, Not Keyword Blocking

Building layered input, action, and output controls that remain testable when language and threats change.

Threat Modeling an AI Application End to End

Assets, trust boundaries, model-specific attacks, tool abuse, data leakage, and concrete mitigations for deployed AI systems.

Prompt Injection Defense in Depth

Why delimiters are not a sandbox, how indirect injection reaches agents, and which architectural controls actually reduce impact.

Apply the category

Put this against a real system.

If a decision in Evaluation and Security is in front of you right now, the fastest version of this is a call: bring the architecture, the failure you are seeing, and the constraint you cannot move.

Discuss the system