The numbers tell a clear story
The data on AI agent trust is moving in the wrong direction. Gravitee's State of AI Agent Security 2026 report found that 88% of organisations reported at least one AI agent security incident in the preceding year.
That is not a sampling anomaly, it reflects the reality that autonomous systems operating at scale will inevitably encounter edge cases their designers did not anticipate, and when they do, the consequences compound quickly.
Developer confidence is eroding in parallel. Only 29% of developers trust the accuracy of AI-generated output, down from 40% the prior year (2025 Stack Overflow Developer Survey). This is a remarkable trajectory for a technology that is simultaneously being adopted at record pace. Developers are using AI tools more, while trusting them less. That tension is unsustainable.
Samsung learned this the hard way in early 2023, when employees inadvertently leaked proprietary semiconductor data through an AI coding assistant. The assistant worked exactly as designed: it accepted the input, processed it, and sent it to external servers for inference. The data was potentially incorporated into the model's training pipeline. The failure was not in the AI.
It was in the absence of governance around what the AI was permitted to see, do, and transmit.
These are not isolated incidents. They are structural symptoms of a fundamental architectural gap: AI systems that can act but cannot prove what they did.