AI Evaluation Gap Found
Enterprises are granting AI agents more autonomy while trusting evaluations less, leading to a reality-alignment problem. Many have shipped agents to production that later failed, highlighting the need for more reliable evaluation methods.

### Introduction to the AI Evaluation Gap A recent study by VentureBeat AI has uncovered a concerning trend in enterprise AI organizations. Despite the increasing autonomy given to AI agents, there is a growing distrust in the evaluations meant to ensure their reliability.
### The Reality-Alignment Problem The research, which involved 157 enterprises, shows that half of the organizations have already deployed AI agents that passed internal evaluations but failed in real-world customer scenarios. This discrepancy suggests a significant reality-alignment problem rather than a coverage issue.
### Lack of Trust in Automated Evaluations Only one in twenty enterprises fully trusts automated evaluation today, with the most cited weakness being the lack of alignment between evaluations and real-world outcomes. Despite this, two-thirds of the organizations are either already deploying agent changes to production based solely on automated evaluation or are working towards this capability, removing human oversight from the process.
Read also

US Invests $5B in AI Science
The US government has launched the Genesis Mission, a $5 billion initiative to fund AI-driven science projects. This effort aims to drive innovation and advancement in the field of AI.

TikTok Tests AI Likeness Tool
TikTok has begun testing an AI likeness detection tool to identify and report AI-generated content. The tool is initially being tested with some US creators.

AI Compute Gap Widens
Enterprises are accelerating AI infrastructure spending without fully understanding the costs, leading to a compute gap. Most organizations lack visibility into their AI compute economics.