Skip to content
Putting technology to work.
Insights to guide decisions and action.

Search articles

AI scores under 50% on ITBench-AA ─ Human-AI collaborative IT operations for clients in 2026

Table of contents · 11 items

On May 27, 2026, the Hugging Face Blog (co-authored by Artificial Analysis and IBM Research) published ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks. ITBench-AA is the first comprehensive benchmark to test 116 enterprise IT tasks—spanning SRE, incident response, network diagnostics, FinOps, and compliance—by having them executed in agentic environments. The results showed GPT-5.5 at 49%, Claude Opus 4.X at 47%, Gemini 3.5 Pro at 41%, DeepSeek R3 at 38%, and Llama 4 at 32%—unveiling the reality that even frontier models cannot complete half of the operational tasks. Around the same time, MIT Technology Review successively published Rethinking organizational design in the age of agentic AI and A reality check on the AI jobs hysteria, making it a shared industry consensus that last year's management initiative to "halve the IT department with AI" was based on overinflated expectations.

From the perspective of supporting IT operations, IT teams, and SRE for mid-sized enterprises through custom development, this means we have entered a phase where organizations abandon "delegating everything wholesale to AI agents" and instead design pair operations based on "humans making judgments while AI handles preparation" as the new mainstream model. Connecting this with the enterprise-wide AI rollout in non-IT industries in Custom Hyatt × ChatGPT Enterprise Development (GH Media), the engineering support in Grab Multi-Agent Internal Help Desk (GH Media), and the AI ROI evaluation in Custom Microsoft AI Cost vs. Headcount Development (GH Media), we organize "human-AI cooperative IT operations" as a custom architectural methodology.

Why human-AI collaborative IT operations are a turning point

DimensionFully automated AI illusion (up to 2025)Human-AI collaborative operations (2026 standard)
Design premiseAI automates 80–90%AI handles 40–50%; humans handle judgment / verification
Demarcation of responsibilityAmbiguous (delegating everything to AI)Explicit (by task type + severity)
Incident responseAI makes autonomous decisionsAI suggests candidates + human approves
Knowledge managementPrompts onlyKnowledge + prompts + runbooks
Training"Leave it to AI""Solving together with AI" skills
KPIAutomation rate+ Error avoidance rate / learning speed
Remedy in case of failureUnclearExplicitly stated in contract
Target scopeAll operationsSelected based on task-specific suitability scores

In other words, human-AI collaborative operations represent a structural shift toward pragmatism: "recognizing what AI cannot do and driving productivity through human-AI pairing."

Three structural changes beneficial to custom development projects

Structure 1: From "across-the-board AI adoption" to "task suitability scores"

In 2024–2025, IT teams at mid-sized enterprises purchased multiple AI products under the promise of "saving labor with AI." However, as ITBench-AA demonstrates, AI suitability varies significantly across tasks. In custom development, we perform a classification of 116 tasks mapped against client operations and deliver "agentization strictly for tasks scoring 70 or higher in AI suitability." This represents a task-granularity version of the ROI evaluation addressed in Custom Microsoft AI Cost vs. Headcount Development (GH Media).

Structure 2: From "autonomous AI execution" to "human-in-the-loop"

For operations that are high-impact and irreversible, such as incident response and change management, human-in-the-loop is essential: AI proposes candidates → human approves → AI executes. Through custom development, we deliver a pair operations infrastructure that includes approval UIs, audit logs, and rollbacks. This is the approval-gated version of the multi-agent design discussed in Grab Multi-Agent Internal Helpdesk (GH Media).

Structure 3: From "training that relies on AI" to "AI collaborative skills"

The ITBench-AA results also demonstrated that "operators who trust AI blindly will fail." Through custom development, we provide training programs and evaluations for new IT skills: questioning AI outputs, designing prompts, and validating results. This is the IT department skill version of the non-IT industry rollout discussed in Hyatt × ChatGPT Enterprise Custom Development (GH Media).

The 5 phases of human-AI collaborative IT operations delivered via custom development

Phase 1: Current state assessment (2–3 weeks)

  • IT operational inventory (116 tasks × internal operations)
  • Inventory of existing AI tools / agents
  • AI suitability scoring (automated / semi-automated / manual)
  • Demarcation of responsibilities for incident / change management
  • Skills map evaluation
  • Risk and ROI matrix

Phase 2: Design (2–3 weeks)

  • Operational models by task (automated / human-in-the-loop / manual)
  • Approval gates + rollback design
  • Audit logs + prompt retention
  • Training program + evaluation system
  • KPIs (automation rate + error avoidance rate + learning speed)
  • Incident runbooks

Phase 3: Implementation (4–5 weeks)

  • Agent foundation (Claude Code / Codex / Copilot / custom in-house)
  • Approval UI (Slack / Teams / custom web UI)
  • Knowledge base (Notion / Confluence / custom in-house RAG)
  • Audit logging (OpenTelemetry + SIEM)
  • Rollback foundation (IaC / database snapshots)
  • Dashboards (Grafana / Datadog)

Phase 4: Pilot rollout (3–4 weeks)

  • Launch operations for tasks with a suitability score of 70 or higher
  • Testing human-in-the-loop user flows
  • KPI measurement + refinement
  • Conducting training programs
  • Feedback loop

Phase 5: Monthly operational reviews (ongoing)

  • Automation rate / error rate by task
  • New model evaluation (official ITBench-AA + custom internal testing)
  • Skill evaluation + career paths
  • Incident root cause analysis
  • Semiannual suitability score updates

Standard technology stack set for custom development

LayerRecommended technologyAlternative
AgentClaude Code / Codex / Copilot WorkspacexAI Skills / custom in-house
Approval UISlack Workflows / Teams Bot / custom web UIDiscord
KnowledgeNotion / Confluence / GitHub WikiCustom in-house RAG
Audit LoggingOpenTelemetry + SIEMDatadog Logs
IaCTerraform / OpenTofu / PulumiAnsible
RollbacksDB snapshot / IaC drift detectionVelero
Evaluationpromptfoo / Langfuse Evals + ITBenchIn-house + manual
DashboardGrafana / Datadog / LookerNew Relic

Which projects need this and which do not

Projects requiring thisProjects not requiring this
IT department has already purchased redundant AI toolsAI not yet adopted
Wants to delegate incident / change management to AIFully automated batch processing
Audit requirements (ISO 27001 / SOC 2 / J-SOX)Not subject to auditing
AI ROI fell short of expectations and requires redesignROI already achieved
Upskilling the IT department is an urgent prioritySufficient existing skills

Six clauses to include in client contracts

ClauseDetailsWhat the client should verify
Target tasksAutomated / human-in-the-loop / manualHandling out-of-scope tasks
Approval gateApprovers + SLAs by taskEscalation thresholds
Remedy in case of failureDemarcation of responsibilities for AI-caused vs. human-caused issuesLitigation / dispute contingency plans
Audit log retentionRetention period + encryption + access controlRegulatory requirements
Handover Upon Project CompletionConfiguration / prompts / runbooks / trainingInternal operational continuity
Incident operations24/7 / mandatory human-in-the-loopEmergency exceptions

Client ROI estimate (assuming an IT department of 60 members and 200 monthly incidents)

ItemCurrent state (across-the-board AI adoption requiring redesign)After implementing collaborative operationsDifference
Average incident resolution time (MTTR)4.5 hours2.5 hours-2 hours / incident
AI-caused incidents (monthly)4 incidents0.5 incidents-3.5 incidents
Training investment effectUnknownVisualized through evaluation systemMeasurable via KPIs
AI tool subscription costs4 million yen / month2.4 million yen / month-¥1,600,000
Staff turnover rateHigh (burnout)Low (sense of learning and growth)Talent retention
Annual benefitEquivalent to approx. 48 million yen + talent retention + explainable ROI

At an hourly rate of 8,000 yen, this translates to an annual workload reduction of 38 million yen + 19.2 million yen in AI tool cost savings. Compared to the substantial cost reductions achieved, the expenses required to establish this operational structure typically prove to be well worth the investment.

Five common pitfalls

Pitfall 1: Mapping all internal operations strictly to the 116 tasks

ITBench-AA is a general-purpose benchmark, and industry-specific tasks require separate evaluation. Always conduct an internal, customized suitability scoring.

Pitfall 2: Oversimplifying human-in-the-loop as "humans just approving"

It is not uncommon for approvers to rubber-stamp everything with "OK," making the process a formality. Institutionalize sampling reviews + randomized testing.

Pitfall 3: Ending training programs with just "how-to tutorials"

Merely learning "how to use Claude" will not build the skill of scrutinizing AI outputs. Include evaluations, failure case sharing, and prompt reviews.

Pitfall 4: Introducing too many AI tools

Adopting everythingClaude Code + Copilot + Cursor + custom tools—leads to operational breakdown. Narrow down to 2–3 products segmented by role and abstract them through a gateway.

Pitfall 5: Leaving liability terms vague for failures

When an AI-caused incident occurs, writing it off as "the fault of the operator who used the AI" will result in staff refusing to use AI. Clearly define the demarcation of responsibilities in contracts and internal policies.

90-day action plan

WeekAction
Week 1〜3Operational inventory + suitability scoring + ROI evaluation
Week 4〜5Operational model design + approval UI design + training program
Week 6〜10Agent foundation + approval UI + audit logging setup
Week 11〜12Pilot task rollout + training execution + KPI measurement
Week 12Company-wide rollout + runbook preparation
Week 13Initial monthly review + ROI dashboard

Summary — Evolving enterprise IT operations: from delegating everything to AI to human-AI pair operations

The findings from ITBench-AA showing frontier models scoring below 50% have scientifically refuted last year's management initiative to "halve the IT department with AI." From the standpoint of supporting mid-sized enterprises' IT operations through custom development, designing task suitability scores, human-in-the-loop, training, approval UIs, and auditing as an integrated whole will become the standard approach going forward.

Because required structures vary significantly depending on IT department size, existing AI tool configurations, and audit requirements, we quote our support scope and fees on an individual basis. If you are facing challenges such as "we purchased multiple AI tools but cannot see the ROI," "AI-related incidents have increased," or "upskilling our IT team is an urgent priority," please feel free to contact us via our contact form.

Sources

Share this articleXFacebook
Kakeru Suzuki

Fascinated by the possibilities of technology, has had a deep interest in programming and digital art since student days

Turn this article's theme into your company's next step

Thinking together, starting from the work you entrust to AI.

We organize your current operations and data to define the scope entrusted to AI, what humans should review, and how to run trials.

  • Target operations
  • Data to use
  • How to verify effectiveness
Consult on AI adoption for your business

You can consult with us from the initial conceptual stage. Details from this article will be carried over to the inquiry form.

Receive the latest articles by email