A Declarative Multi-Agent Framework for Reducing On-Call Operational Burden

Twinkle Kariya
2026

Abstract

Large-scale cloud-native systems generate continuous streams of operational alerts across distributed microservice architectures. On-call engineers must manually triage these alerts by correlating signals from heterogeneous observability tools, a process that is time-consuming, cognitively demanding, and prone to error. Despite advances in monitoring and anomaly detection, incident triage remains largely manual. This paper presents a declarative, large language model (LLM)–driven multi-agent approach to automating incident triage and Service Level Objective (SLO) monitoring. The proposed design constrains agent behavior using domain-expertauthored investigation workflows, enabling deterministic execution and reproducibility while preserving operational safety. The framework integrates a unified tool execution layer for interacting with diverse observability systems and an enhanced retrieval-augmented generation (RAG) pipeline optimized for operational knowledge retrieval. The approach has been evaluated in a production cloud environment spanning multiple microservices and geographic regions. Results show reductions in high-severity incident triage time from approximately 30 minutes to under 5 minutes, alert acknowledgement latency from minutes to seconds, and service onboarding effort from weeks to days. These findings suggest that constrained multi-agent systems can substantially reduce on-call cognitive load while maintaining reliability and human oversight.

Research Areas

×