Recorded signals
number
End-to-end or provider-call duration, depending on the event type.
object
Prompt, completion, and total token counts when the provider returns usage.
number
Estimated provider cost based on the configured model registry; treat this as an estimate.
boolean
required
Whether the operation completed successfully.
object
Provider, model, tool, channel, run, and other structured diagnostic context.
Three levels of diagnosis
Recommended alerts
- AI queue age or failed jobs exceed normal bounds.
- Provider success rate or time to first token degrades.
- Tool error rate spikes by tool name.
- Fallback and human-escalation rate changes unexpectedly.
- Qdrant or embedding calls fail or return no context unusually often.
- Observability/analytics workers repeatedly fail to flush.
