A VentureBeat Pulse Research survey of 157 technical leaders at organizations with 100+ employees reveals that half (50%) have deployed an AI agent or LLM feature in the past year that passed internal evaluations but caused a customer-facing failure, with a quarter experiencing this more than once. Despite this, only 5% fully trust automated evaluation today, and the most-cited limitation (29%) is that evaluations align poorly with real-world outcomes.
Two-thirds of organizations (66%) already permit or are actively engineering toward fully automated, zero-human-in-the-loop deployment: 34% allow it for low-risk agents now, while 33% are building pipelines to enable it within twelve months. The evaluation tools meant to govern this autonomy remain fragmented and immature, with the most common primary tools being model providers' native evals or no dedicated tooling (17% each), and only about a quarter running real-time quality checks on live production traffic.
The research identifies an "evaluation gap" — the growing distance between the autonomy enterprises are granting agents and the trust they place in the evaluations meant to govern that autonomy. The core finding: the autonomy is arriving faster than the assurance, with organizations discovering that a passing eval is not the same as a working agent in production.