Semantic checks
Mibo's AI-powered checks compare response meaning with criteria you describe, such as grounded, complete, safe, or aligned. They score the supplied response; they do not establish truth or guarantee quality on their own.
Mibo combines active evaluation before release with passive evaluation of real traces after launch, so teams can inspect semantic quality, procedural evidence, and the limits of each signal.
Mibo turns expected agent behavior into evidence you can inspect instead of treating a successful HTTP response as proof that the interaction was reliable.
Mibo's AI-powered checks compare response meaning with criteria you describe, such as grounded, complete, safe, or aligned. They score the supplied response; they do not establish truth or guarantee quality on their own.
Mibo's deterministic checks inspect observable facts such as node or tool calls, arguments, attributes, HTTP status, schema, and response time. A missing trace field is reported as missing instrumentation rather than silently passing.
After a real interaction, your system can send a trace for asynchronous evaluation against applicable active tests. Passive results cover what you send and what your tests can observe—not every possible conversation.
Use the same expectations in a controlled active run and in passive traces, while keeping their evidence and limits visible.
01
Describe the scenario and assertions that matter, including semantic criteria and observable execution facts such as a required tool call or status.
02
Choose a scenario, send it to your connected Agent, and review the response with its execution evidence before release. Active runs do not prove unobserved cases.
03
Send a canonical trace after a real interaction. Mibo stores it and evaluates it asynchronously against applicable active tests outside the customer request path; missing fields can limit the result.
Use plain-language criteria to evaluate meaning, safety, completeness, and alignment. The result depends on the response and context you provide, so it is evidence for investigation—not an automatic quality guarantee.
Inspect tool or node calls, arguments, attributes, HTTP status, schema, and timing with deterministic checks. These signals are limited by the instrumentation and fields present in the active response or passive trace.
Mibo evaluates the behavior you describe, including response meaning and observable execution evidence such as routes, tool calls, values, status, schema, and timing. Semantic checks assess criteria in plain language; procedural checks verify recorded facts. Neither proves every possible interaction or guarantees overall quality.
Yes. Active evaluation sends selected scenarios to a connected staging or production Agent so you can review the response and execution evidence. You can reuse the same test definitions for passive evaluation when real interactions arrive as traces.
Your system sends a trace after the interaction. Mibo stores it, matches applicable active tests, and evaluates it asynchronously in the background, so Mibo does not add another request to the customer path.
No. Mibo evaluates configured checks against the responses and traces available to it. Narrow criteria, unobserved edge cases, unavailable classification, and missing instrumentation can limit the signal, so active runs and passive traces should be reviewed together.
Use the maintained documentation to define scenarios, inputs, assertions, and passive behavior.
Read test creation docs →Define scenarios, run active checks before release, and reuse them for passive production evaluation.
Explore AI agent testing →Turn real agent expectations into concrete scenarios with semantic and procedural assertions you can run and inspect.
Explore AI agent test cases →Test safety-related behavior before release and inspect bounded evidence from real interactions.
Explore AI agent safety testing →Separate observed scenario traffic from idle and unmatched cases without treating coverage as proof of quality.
Explore test coverage →Connect deliberate checks before release with evaluation of available production traces after real interactions.
Explore regression testing →Follow a practical workflow for test cases, active runs, passive traces, and interpretation limits.
Read the guide →Evaluate the behavior users experience in production through real traces and investigate regressions.
Explore monitoring →Choose behavioral and response metrics without hiding their interpretation limits.
Explore metrics →Run the canonical live reliability check with a synthetic customer request.
Run a live check →