On May 27, 2026, the Hugging Face Blog (co-authored by Artificial Analysis and IBM Research) published ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks. ITBench-AA is the first comprehensive benchmark to test 116 enterprise IT tasks—spanning SRE, incident response, network diagnostics, FinOps, and compliance—by having them executed in agentic environments. The results showed GPT-5.5 at 49%, Claude Opus 4.X at 47%, Gemini 3.5 Pro at 41%, DeepSeek R3 at 38%, and Llama 4 at 32%—unveiling the reality that even frontier models cannot complete half of the operational tasks. Around the same time, MIT Technology Review successively published Rethinking organizational design in the age of agentic AI and A reality check on the AI jobs hysteria, making it a shared industry consensus that last year's management initiative to "halve the IT department with AI" was based on overinflated expectations.
From the perspective of supporting IT operations, IT teams, and SRE for mid-sized enterprises through custom development, this means we have entered a phase where organizations abandon "delegating everything wholesale to AI agents" and instead design pair operations based on "humans making judgments while AI handles preparation" as the new mainstream model. Connecting this with the enterprise-wide AI rollout in non-IT industries in Custom Hyatt × ChatGPT Enterprise Development (GH Media), the engineering support in Grab Multi-Agent Internal Help Desk (GH Media), and the AI ROI evaluation in Custom Microsoft AI Cost vs. Headcount Development (GH Media), we organize "human-AI cooperative IT operations" as a custom architectural methodology.
Why human-AI collaborative IT operations are a turning point
| Dimension | Fully automated AI illusion (up to 2025) | Human-AI collaborative operations (2026 standard) |
|---|---|---|
| Design premise | AI automates 80–90% | AI handles 40–50%; humans handle judgment / verification |
| Demarcation of responsibility | Ambiguous (delegating everything to AI) | Explicit (by task type + severity) |
| Incident response | AI makes autonomous decisions | AI suggests candidates + human approves |
| Knowledge management | Prompts only | Knowledge + prompts + runbooks |
| Training | "Leave it to AI" | "Solving together with AI" skills |
| KPI | Automation rate | + Error avoidance rate / learning speed |
| Remedy in case of failure | Unclear | Explicitly stated in contract |
| Target scope | All operations | Selected based on task-specific suitability scores |
In other words, human-AI collaborative operations represent a structural shift toward pragmatism: "recognizing what AI cannot do and driving productivity through human-AI pairing."
Three structural changes beneficial to custom development projects
Structure 1: From "across-the-board AI adoption" to "task suitability scores"
In 2024–2025, IT teams at mid-sized enterprises purchased multiple AI products under the promise of "saving labor with AI." However, as ITBench-AA demonstrates, AI suitability varies significantly across tasks. In custom development, we perform a classification of 116 tasks mapped against client operations and deliver "agentization strictly for tasks scoring 70 or higher in AI suitability." This represents a task-granularity version of the ROI evaluation addressed in Custom Microsoft AI Cost vs. Headcount Development (GH Media).
Structure 2: From "autonomous AI execution" to "human-in-the-loop"
For operations that are high-impact and irreversible, such as incident response and change management, human-in-the-loop is essential: AI proposes candidates → human approves → AI executes. Through custom development, we deliver a pair operations infrastructure that includes approval UIs, audit logs, and rollbacks. This is the approval-gated version of the multi-agent design discussed in Grab Multi-Agent Internal Helpdesk (GH Media).
Structure 3: From "training that relies on AI" to "AI collaborative skills"
The ITBench-AA results also demonstrated that "operators who trust AI blindly will fail." Through custom development, we provide training programs and evaluations for new IT skills: questioning AI outputs, designing prompts, and validating results. This is the IT department skill version of the non-IT industry rollout discussed in Hyatt × ChatGPT Enterprise Custom Development (GH Media).
The 5 phases of human-AI collaborative IT operations delivered via custom development
Phase 1: Current state assessment (2–3 weeks)
- IT operational inventory (116 tasks × internal operations)
- Inventory of existing AI tools / agents
- AI suitability scoring (automated / semi-automated / manual)
- Demarcation of responsibilities for incident / change management
- Skills map evaluation
- Risk and ROI matrix
Phase 2: Design (2–3 weeks)
- Operational models by task (automated / human-in-the-loop / manual)
- Approval gates + rollback design
- Audit logs + prompt retention
- Training program + evaluation system
- KPIs (automation rate + error avoidance rate + learning speed)
- Incident runbooks
Phase 3: Implementation (4–5 weeks)
- Agent foundation (Claude Code / Codex / Copilot / custom in-house)
- Approval UI (Slack / Teams / custom web UI)
- Knowledge base (Notion / Confluence / custom in-house RAG)
- Audit logging (OpenTelemetry + SIEM)
- Rollback foundation (IaC / database snapshots)
- Dashboards (Grafana / Datadog)
Phase 4: Pilot rollout (3–4 weeks)
- Launch operations for tasks with a suitability score of 70 or higher
- Testing human-in-the-loop user flows
- KPI measurement + refinement
- Conducting training programs
- Feedback loop
Phase 5: Monthly operational reviews (ongoing)
- Automation rate / error rate by task
- New model evaluation (official ITBench-AA + custom internal testing)
- Skill evaluation + career paths
- Incident root cause analysis
- Semiannual suitability score updates
Standard technology stack set for custom development
| Layer | Recommended technology | Alternative |
|---|---|---|
| Agent | Claude Code / Codex / Copilot Workspace | xAI Skills / custom in-house |
| Approval UI | Slack Workflows / Teams Bot / custom web UI | Discord |
| Knowledge | Notion / Confluence / GitHub Wiki | Custom in-house RAG |
| Audit Logging | OpenTelemetry + SIEM | Datadog Logs |
| IaC | Terraform / OpenTofu / Pulumi | Ansible |
| Rollbacks | DB snapshot / IaC drift detection | Velero |
| Evaluation | promptfoo / Langfuse Evals + ITBench | In-house + manual |
| Dashboard | Grafana / Datadog / Looker | New Relic |
Which projects need this and which do not
| Projects requiring this | Projects not requiring this |
|---|---|
| IT department has already purchased redundant AI tools | AI not yet adopted |
| Wants to delegate incident / change management to AI | Fully automated batch processing |
| Audit requirements (ISO 27001 / SOC 2 / J-SOX) | Not subject to auditing |
| AI ROI fell short of expectations and requires redesign | ROI already achieved |
| Upskilling the IT department is an urgent priority | Sufficient existing skills |
Six clauses to include in client contracts
| Clause | Details | What the client should verify |
|---|---|---|
| Target tasks | Automated / human-in-the-loop / manual | Handling out-of-scope tasks |
| Approval gate | Approvers + SLAs by task | Escalation thresholds |
| Remedy in case of failure | Demarcation of responsibilities for AI-caused vs. human-caused issues | Litigation / dispute contingency plans |
| Audit log retention | Retention period + encryption + access control | Regulatory requirements |
| Handover Upon Project Completion | Configuration / prompts / runbooks / training | Internal operational continuity |
| Incident operations | 24/7 / mandatory human-in-the-loop | Emergency exceptions |
Client ROI estimate (assuming an IT department of 60 members and 200 monthly incidents)
| Item | Current state (across-the-board AI adoption requiring redesign) | After implementing collaborative operations | Difference |
|---|---|---|---|
| Average incident resolution time (MTTR) | 4.5 hours | 2.5 hours | -2 hours / incident |
| AI-caused incidents (monthly) | 4 incidents | 0.5 incidents | -3.5 incidents |
| Training investment effect | Unknown | Visualized through evaluation system | Measurable via KPIs |
| AI tool subscription costs | 4 million yen / month | 2.4 million yen / month | -¥1,600,000 |
| Staff turnover rate | High (burnout) | Low (sense of learning and growth) | Talent retention |
| Annual benefit | — | — | Equivalent to approx. 48 million yen + talent retention + explainable ROI |
At an hourly rate of 8,000 yen, this translates to an annual workload reduction of 38 million yen + 19.2 million yen in AI tool cost savings. Compared to the substantial cost reductions achieved, the expenses required to establish this operational structure typically prove to be well worth the investment.
Five common pitfalls
Pitfall 1: Mapping all internal operations strictly to the 116 tasks
ITBench-AA is a general-purpose benchmark, and industry-specific tasks require separate evaluation. Always conduct an internal, customized suitability scoring.
Pitfall 2: Oversimplifying human-in-the-loop as "humans just approving"
It is not uncommon for approvers to rubber-stamp everything with "OK," making the process a formality. Institutionalize sampling reviews + randomized testing.
Pitfall 3: Ending training programs with just "how-to tutorials"
Merely learning "how to use Claude" will not build the skill of scrutinizing AI outputs. Include evaluations, failure case sharing, and prompt reviews.
Pitfall 4: Introducing too many AI tools
Adopting everything—Claude Code + Copilot + Cursor + custom tools—leads to operational breakdown. Narrow down to 2–3 products segmented by role and abstract them through a gateway.
Pitfall 5: Leaving liability terms vague for failures
When an AI-caused incident occurs, writing it off as "the fault of the operator who used the AI" will result in staff refusing to use AI. Clearly define the demarcation of responsibilities in contracts and internal policies.
90-day action plan
| Week | Action |
|---|---|
| Week 1〜3 | Operational inventory + suitability scoring + ROI evaluation |
| Week 4〜5 | Operational model design + approval UI design + training program |
| Week 6〜10 | Agent foundation + approval UI + audit logging setup |
| Week 11〜12 | Pilot task rollout + training execution + KPI measurement |
| Week 12 | Company-wide rollout + runbook preparation |
| Week 13 | Initial monthly review + ROI dashboard |
Summary — Evolving enterprise IT operations: from delegating everything to AI to human-AI pair operations
The findings from ITBench-AA showing frontier models scoring below 50% have scientifically refuted last year's management initiative to "halve the IT department with AI." From the standpoint of supporting mid-sized enterprises' IT operations through custom development, designing task suitability scores, human-in-the-loop, training, approval UIs, and auditing as an integrated whole will become the standard approach going forward.
Because required structures vary significantly depending on IT department size, existing AI tool configurations, and audit requirements, we quote our support scope and fees on an individual basis. If you are facing challenges such as "we purchased multiple AI tools but cannot see the ROI," "AI-related incidents have increased," or "upskilling our IT team is an urgent priority," please feel free to contact us via our contact form.
Sources
- ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks(Hugging Face Blog 2026-05-27)
- Rethinking organizational design in the age of agentic AI(MIT Technology Review 2026-05-26)
- A reality check on the AI jobs hysteria(MIT Technology Review 2026-05-26)
- Hyatt × ChatGPT Enterprise Client Project (GH Media)
- Grab Multi-Agent Internal Helpdesk (GH Media)
- Microsoft AI Costs vs. Headcount Custom Development (GH Media)









