Observability & evaluation for AI

See what your AI is doing. Make it better.

Follow every trace, find what went wrong, and turn real examples into evaluations. One place to understand and improve your AI workflows.

From a single request to a clearer picture.

Follow every trace

Inspect inputs, outputs, spans, latency, and scores in the context of the whole workflow.

Evaluate real examples

Build datasets from the cases that matter. Test changes with repeatable evaluation runs.

Make quality visible

Use scorers, human reviews, and dashboards to understand how your application performs.

Start with the evidence you already have. Trace a workflow, review an unexpected result, and keep that example as a test for the next change.

A few things you might be wondering.

What is Datool?

Datool helps teams inspect AI traces, build evaluation datasets, and measure quality with scorers. Follow a request from input to output and see where it needs work.

What can I evaluate?

Use recorded traces or dataset cases to evaluate your AI workflows. Combine code-based and LLM scorers with human reviews, then compare results across evaluation runs.

How do I get started?

Sign in with your approved Google account, select a workspace and project, and follow the project setup instructions to connect your application.

Your next improvement starts with a trace.

Connect your application and take a closer look.