ATX AI Eval — Universal Evaluation Suite
ATX AI Eval is a universal evaluation suite for AI agents and LLM applications, available on its own or inside ATX Guardrails. It grades agent behaviour with a multi-judge LLM jury, attaches to any framework through OpenTelemetry instrumentation, and generates compliance-ready documentation — so teams move agents into production on evidence rather than guesswork.
Six metrics are live today, with eight agent-native metrics landing next — CLEAR, agent handoff quality, context drift and MCP security among them. The same harness runs as a runtime guardrail, a pytest suite and a CI/CD gate, and deployment spans SaaS, local Docker, in-VPC and fully air-gapped.