The Most Dangerous AI Agent Failure Might Return 200 OK
An agent processes a refund. Its credentials are valid, it calls the correct API, authentication and authorization both pass, and the service returns 200 OK. It also refunded the wrong customer. Every dashboard stays green because the stack measured whether the request succeeded, not whether it was the right request. This month's lead story walks you through the variants: an agent retrying after a lost acknowledgment and refunding twice, or completing five tool calls while skipping the compliance export that policy required first.
The proposed fix moves correctness out of the prompt. Give consequential tools deterministic preconditions and postconditions so the orchestration layer can do the checking: the model reasons, and the system verifies. Carry an intent_id alongside the usual request and trace IDs, so a timeout can be answered by asking whether that objective already produced a refund. Make mutations idempotent. And never let the agent grade its own homework. If it misread the intent once, asking it to confirm its own work reproduces the mistake.