Observability
How to monitor AI application latency
Track model, retrieval and tool-call latency separately.
Problem
AI-generated code often optimizes for getting the happy path working. Production systems also need explicit behavior for failures, limits, security boundaries, and operations.
Why it matters
This pattern reduces avoidable outages and makes system behavior easier to understand when dependencies become slow, unavailable, or unpredictable.
Production pattern
measure retrieval_ms, model_ms, tool_ms, total_msCommon mistakes
- Assuming default framework behavior is production-safe.
- Ignoring failure and recovery paths.
- Logging sensitive values while debugging.
- Adding complexity without monitoring it.
AfterCode checklist
- Define expected behavior.
- Test success and failure cases.
- Add observable signals.
- Document operational ownership.
Next step
Run the broader production checklist to find adjacent risks around this pattern.
Open production checklist