Reports that late-night alerts now recover automatically are always welcome. On-call pages drop, next-day productivity rises, and operational cost visibility improves. Up to this point, everything goes according to the proposal.
The problem emerges after that. About six months in, an incident occurs of a kind that AI cannot fix. It spans multiple causes, and the symptoms do not match anything in the past. Only then do the responders realize: for the past six months, they have not properly read the logs of this system.
This is not a matter of competence. It is a structural problem: the opportunity for practice has disappeared.
What was pointed out in 1983
This phenomenon has a name. It was identified in the 1983 paper "Ironies of Automation" by ergonomics researcher Lisanne Bainbridge.
The core point is this: automation takes over simpler tasks first. Consequently, only the difficult tasks that cannot be automated remain for humans. However, humans build the intuition needed for difficult tasks through the repetition of simpler tasks. When the simpler tasks are taken away, preparedness for the difficult tasks is lost as well.
Translating this into systems operations: routine incidents—such as full disks, crashed processes, or expired certificates—used to be opportunities for staff to safely build an intuition for "how this system breaks." AI is taking over those exact opportunities first.
The aviation industry's answer: "Make training mandatory"
The aviation industry faced this exact problem earlier. As autopilots became widespread, opportunities for manual flying decreased, yet crews still needed to be prepared for situations that autopilot could not handle.
Under US regulations, commercial flight crews are required to undergo recurrent training every six months. The training includes scenarios like an engine failure during takeoff—events you never want to experience in a real aircraft, but must be able to handle if they occur.
The order of priorities is critical. Training is not treated as "something to do if there is time"; it is made a condition for maintaining qualifications. If it is structured as something done only when there is time, it gets skipped during the busiest months first.
Similar initiatives have begun in software operations. Incident management platform Rootly has partnered with Uptime Labs to provide simulated drill environments mimicking e-commerce site outages. Participating engineers act as incident commanders using real monitoring tools, required to make decisions and report to stakeholders. Because LLMs play the roles of stakeholders, drills can be run without gathering people.

What clients should include in contracts
When outsourcing maintenance and introducing AI-driven automated response, certain line items do not appear in cost estimates. These are also items that are easy to cut because nothing breaks immediately when they are removed.
1. Post-incident sharing for AI-resolved issues. This refers to whether details of automatically resolved incidents appear in monthly reports. A raw figure like "47 automatic recoveries" tells no one how the system is breaking. It needs to remain recorded in a readable form detailing what broke, how it broke, and how it was restored. If you have not aligned with your contractor on what is defined as normal, this is the first area that becomes vague.
2. Responsibility for runbook updates. To the extent that AI handles issues, runbooks meant for humans stop being updated. When an incident that AI cannot handle occurs a year later, you want to avoid opening a runbook that is two years out of date. Define who updates it and at what frequency.
3. Frequency of human-led drills. Even once a year is fine. It comes down to whether the contract contains a commitment to deliberately set aside time to "perform initial triage without relying on AI." If this is left blank, it will not happen.
| Item | Treatment in estimates | What happens when cut |
|---|---|---|
| Post-sharing of automated responses | Treated as included under "Reports" | Incident trends become invisible to everyone |
| Runbook updates | Included in initial setup, none thereafter | Outdated procedures surface during emergencies |
| Human-led drills | No line item to begin with | No capable personnel remain internally or externally |
For the third line, it is practical to add it when reviewing the working hours and scope of your maintenance contract. It is much easier to get approved than negotiating it from scratch in a brand-new contract.
This does not mean you should avoid AI
To be clear, we do not believe that choosing not to introduce automated response is safer. Automating routine incidents protects personnel from burnout. Reducing the number of late-night wake-up calls holds immense value.
The same logic applies to mechanisms that automatically generate vulnerability fixes. Introducing them reduces human labor, but whether that eliminated work included the step of "inspecting and understanding" changes the compensatory measures required afterward.
What needs to be distinguished are the following two:
- Work reduction (spending less time manually executing tasks) translates directly into profit
- Understanding reduction (no longer observing what is happening) results in a bill that comes due later
Checking once every six months whether introducing automated response has remained confined to the former is all it takes to notice that the three items listed above are missing.
What to do next
Ask your contractor to provide the number and details of incidents automatically recovered over the past three months. If they can only provide the count, the first item is already missing at that point.
Review what they provide and verify whether at least one person internally can explain "how this was handled." If no one can explain it, it is time to add a drill item during your next contract renewal.
GleamHub provides consultations on development, AI, and automation covering reviews of maintenance and operations setups, handover design following the adoption of AI automated responses, and the development of incident response runbooks. Because the approach depends on your current operational structure and outsourcing scope, please reach out via our contact form.









