Recreate the decision safely.
Execute a candidate against recorded requests and, where appropriate, recorded tool results or a controlled read environment. Compare structural fields, schemas, semantics, errors, expected cost, and intended side effects.
- Good for coverage across known examples
- Finds deterministic mismatch before a candidate sees live traffic
- Cannot prove behavior against new request shapes or changed external state