Google open-sources AQuA, a quality agent that samples production sessions to diagnose a live agent, at $3.76 for a 32-session sweep
An agent that checks a production agent from outside the request path, reading but never changing it.
// Key points
- AQuA (Ambient Quality Agent) is an open-source reference implementation in the google/adk-recipes repository; it runs in your own Google Cloud project on a schedule, after each deployment or on demand, sampling up to 1,000 sessions per run.
- It works in five stages: sample, review against a nine-point checklist, cluster, verify clusters against full transcripts, and track insights across runs; it never sits in the request path, changes the agent or opens pull requests itself.
- In a travel-concierge example, 32 sessions produced 42 findings in 9 clusters; the verifier rejected 3, leaving 6 issues, and after two one-line prompt fixes full-session passes rose from 5/32 to 13/32.
- Cost: $0.70 for a 96-session single-agent sweep and $3.76 for a 32-session multi-agent sweep, about $0.12 per session.
Builder's takeThe step I value most is verification: have the model find issues, then check them against the raw transcripts to drop false positives, the same reason I insist on primary sources. Conversational products like AI Interview can copy this: sample a few dozen real sessions after each release for a few dollars, far cheaper than waiting for complaints.
// Background · from #AI agents
Full timeline →- Oct 10 Hugging Face shows ML-Intern: an agent trains six models in a few days for about $103 in total
- Oct 8 Local sandboxing for GitHub Copilot is generally available, restricting agent commands’ file, network and credential access with enforceable enterprise policies
- Oct 8 Microsoft makes MXC agent containers generally available on Windows as NVIDIA opens RTX Spark laptop preorders with 1 petaflop of FP4 and up to 128GB of unified memory