AI Agent Reliability Changes When Tools Look Like Code
A 14-model study finds that the interface around an AI agent can materially change tool accuracy, scaling behavior,…
Section
Artificial Intelligence (AI)
From the newsroom
A 14-model study finds that the interface around an AI agent can materially change tool accuracy, scaling behavior,…
A controlled 2,190-video study shows why correct final counts can conceal missed events, incomplete traces, and fragile temporal…
Quantum computing is not the next AI. It could become the specialized engine that expands what AI and…
A new research pipeline tests whether dependency-rich software training data can make long-context AI more reliable across files,…
A USC-led study finds that voice-transcription rewrites can reduce language-model accuracy more than ordinary keyboard mistakes, especially on…
A Microsoft Azure and UT Austin research framework uses formal constraints, simulation, and feedback to search for better…
A new enterprise-document benchmark shows why accurate values alone do not make an AI extraction system dependable: completeness,…
A new benchmark finds that AI evaluators often accept failed computer-use agent runs as successes, creating a measurement…
A new study tests whether modular visual and spatial scaffolding can help general-purpose AI reason about 3D scenes…