Production traces
Send a canonical trace with the interaction, response, and execution details after a real user request through OpenTelemetry or Mibo's Your API path.
Mibo evaluates production traces against the active tests you define, stores asynchronous evidence for investigation, and stays outside the customer request path.
A request can finish successfully while an agent routes to the wrong specialist, calls a tool with unsafe arguments, or gives an incomplete answer. Mibo turns the production trace into evidence you can investigate instead of treating availability as proof of reliable behavior.
Send a canonical trace with the interaction, response, and execution details after a real user request through OpenTelemetry or Mibo's Your API path.
Mibo stores the trace and schedules applicable active tests in the background. Passive evaluation does not send another request to your system.
Review semantic behavior and procedural facts such as tool calls, arguments, status, schema, tokens, and response time when your trace includes the needed fields.
Keep production evaluation separate from the customer interaction while making the evidence and its limits visible to the team responsible for the agent.
01
Your system handles the user request, then sends the canonical trace to Mibo with the input, output, and execution details available to evaluate.
02
Mibo stores the trace and evaluates it asynchronously against applicable active tests. The ingestion response does not wait for the evaluation to finish.
03
Compare results over time and open the evidence behind a failure. Use the findings to prioritize a fix; Mibo does not automatically remediate the agent.
AI-powered semantic assertions compare the response with criteria you describe, such as being grounded, complete, safe, or aligned. They provide evidence for investigation, not a guarantee of truth or overall quality.
Deterministic assertions inspect recorded facts such as node or tool calls, arguments, attributes, HTTP status, schema, token usage, and response time. Missing fields are reported as missing instrumentation rather than a silent pass or fail.
An agent can return a successful HTTP response while choosing the wrong tool, inventing an outcome, or giving an incomplete answer. Mibo evaluates configured behavior and trace evidence, not only whether the endpoint was reachable.
Your system sends a canonical trace after the interaction. Mibo stores it, finds the applicable active tests, and evaluates it asynchronously in the background. It does not call your agent again as part of passive evaluation.
No. Your system handles the customer request first and sends the trace to Mibo separately. The ingestion response is separate from the background evaluation.
No. Mibo evaluates the active tests and trace fields available to it. Unobserved edge cases, narrow criteria, unavailable classification, and missing instrumentation can limit the signal, and Mibo does not automatically remediate the agent.
Trace payloads are encrypted at rest, and product access is scoped to the account that owns them. Send synthetic or appropriately governed data that fits your own privacy requirements.
Review the supported trace paths, asynchronous evaluation flow, passive behavior, and missing-instrumentation limits.
Read passive testing docs →See the span fields and attributes Mibo can use for semantic and procedural evaluation.
Read the trace reference →Understand the semantic and procedural checks behind active runs and passive traces.
Explore evaluation →Define scenarios, run active checks before release, and reuse them for production traces.
Explore AI agent testing →Separate observed scenario traffic from idle and unmatched cases without treating coverage as proof of quality.
Explore test coverage →Connect deliberate checks before release with evaluation of available production traces after real interactions.
Explore regression testing →Run the canonical live reliability check with a synthetic customer request.
Run a live check →Follow a practical workflow for active tests, passive traces, and interpretation limits.
Read the guide →