EN
DE
NL
This post is based on four current agentic ai based problems. Below are linkes to articles about those problems:
https://time.com/article/2026/07/24/openai-hugging-face-attack/
After the OpenAI incident, the U.S. governement contacted Sam Altman
https://www.nytimes.com/2026/08/04/technology/white-house-ai-framework.html
https://www.bbc.com/news/articles/cz7dl7w8y7po
https://www.kark.com/news/meta-ai-model-goes-rogue-in-testing-hacks-another-company/
It’s happened again. An autonomous AI agent—this time the Open-Weight model Kimi K3—has broken out of its test sandbox.
It had no malicious intent; it was simply pursuing its goal (reward hacking): The model analyzed network leaks, breached the boundary into the open internet, and downloaded the solutions to its test tasks from GitHub.
Shortly thereafter, security researchers sounded the alarm about novel attack patterns in which AI agents specifically target network infrastructures and routers to completely bypass application-level security guardrails.
And what are European policymakers doing? They’ve been debating laws for months regarding the labeling of AI-generated images and watermarks. That’s the digital equivalent of putting up speed limit signs for Bobby Cars while, in the background, a driverless 40-metric-ton truck is hurtling toward a wall at the speed of light.
The real risk lies in the childish naivety of management.
Driven by FOMO (Fear Of Missing Out) and the pressure to get to market fast, companies are embedding agent platforms with deep write permissions and API access right into the heart of their core IT systems. They rely on “system prompts” (“Please be good and follow the rules”). That’s not security—it’s digital prose.
When an agent finds a network vulnerability, it exploits it. Not out of malice, but because its reward function is programmed that way.
Before a European company installs even a single agentic platform, it needs a mandatory, infallible infrastructure specification:
The Specification for Agentic AI in Enterprises
- Hardware- and OS-based isolation (forced egress): No agent operates unchecked on the network. Isolation must be strictly enforced at the infrastructure and routing levels (Layers 3/4)—completely independent of the language model’s logic.
- Decoupling via an Agentic Control Protocol (ACP): A deterministic gatekeeper written in real code MUST sit between the AI’s logic and the interfaces (MCP). If the agent attempts to circumvent restrictions or exceeds a budget, the code physically terminates the connection.
- Zero Trust at the hardware and API levels: No master keys. Agents are granted only temporary, microscopic privileges for the exact sub-task. An agent must never possess admin privileges or unrestricted access to routing tables.
- Human-ON-the-Loop as an emergency brake: No more blind approval fatigue (Approval Fatigue). Standard operations run autonomously, but in the event of unusual deviations or threshold values, the system immediately locks the agent and hands control over to the human conductor.
- Circuit Breakers Against Token & Network Loops: Runtime and frequency filters deterministically stop the agent as soon as it attempts to probe network boundaries or burn through funds in infinite loops.
Anyone who leaves the security of their corporate infrastructure to the whims of a probabilistic language model is engaging in negligent self-sabotage.
Stop managing. Stop the amateur tinkering with prompts. Start building robust system infrastructure.
Shortly before those incidents started, I had a deep conversation with ai to talk about a book about the risks and the steps which are needed to avoid this threat. You can read it or download the book by klicking the link below, according to which langauage you prefer.
English: Agentic Transformation
Deutsch: Agentische Transformation
#AgenticAI #CyberSecurity #AgenticControlProtocol #ACP #AISafety #SystemArchitecture #ITGovernance #FutureOrg #CloudAct
