Observability and Fault-Tolerance Techniques for LLM-Agent Pipelines: A Systematic Comparison with Traditional Microservice Architectures
Keywords:
Large Language Model (LLM) Agents, Observability, Fault Tolerance, Microservice Architectures, Distributed Tracing, AI Reliability EngineeringAbstract
The deterministic failure and observability assumptions of blocking concurrent systems, such as microservice systems, impose hard constraints on introducing the non-deterministic failures that Large Language Model (LLM) agents can introduce into enterprise automation pipelines. This paper systematically reviews and compares key techniques, determines which can be adapted, and which should be developed as novel techniques used to support the adoption of LLM agents. The researchers then conducted an extensive literature search and analysed 170 artefacts using a systematic approach: 85 peer-reviewed articles, 30 technical reports, 50 open-source data repositories, and 5 benchmark data sets. The findings showed that by the end of the AI Workflow, distributed tracing captured 92% of microservice failures and 78% of LLM-agent failures, compared with health checks, which dropped from 95% at the start of the AI Workflow. LLM-aware retry mechanisms boosted recovery by 85% over the conventional ones, and checkpointing lowered restart overhead by approximately 30%. The results indicate that traditional observability tools remain valuable but must be complemented with semantic observability, hallucination evaluation, and AI-specific evaluation. They propose a risk-oriented strategy for selecting observability based on the complexity of LLM agents and operational requirements. When deploying agents to autonomous applications in a complex real-world environment, reliability requirements are strict. They should be studied to implement a hybrid approach that combines distributed systems engineering and AI-native practices.
Downloads
References
N. P. Rana, R. Pillai, B. Sivathanu, and N. Malik, “Assessing the nexus of Generative AI adoption, ethical considerations and organizational performance,” Technovation, vol. 135, Art. no. 103064, 2024. https://www.sciencedirect.com/science/article/pii/S0166497224001147
S. W. Jin, N. Wang, L. Zhang, X. Tian, B. Shi, and B. Zhao, “A review of AI-driven automation technologies: Latest taxonomies, existing challenges, and future prospects,” Computers, Materials & Continua, vol. 84, no. 3, pp. 3961–4018, 2025. https://www.sciopen.com/local/article_pdf/10.32604/cmc.2025.067857.pdf
J. Kosińska, B. Baliś, M. Konieczny, M. Malawski, and S. Zieliński, “Toward the observability of cloud-native applications: The overview of the state-of-the-art,” IEEE Access, vol. 11, pp. 73036–73052, 2023. https://ieeexplore.ieee.org/iel7/6287639/10005208/10141603.pdf
I. Tzanettis, C. M. Androna, A. Zafeiropoulos, E. Fotopoulou, and S. Papavassiliou, “Data fusion of observability signals for assisting orchestration of distributed applications,” Sensors, vol. 22, no. 5, Art. no. 2061, 2022. https://www.mdpi.com/1424-8220/22/5/2061
C. Wang, H. Tang, Y. Zhao, W. You, J. Yang, and H. Qiu, “Losing control: Exposing security weaknesses of Kubernetes control plane interfaces,” Computers & Security, Art. no. 104704, 2025. https://www.sciencedirect.com/science/article/pii/S0167404825003931
F. C. Ogenyi, C. N. Ugwu, and O. P.-C. Ugwu, “Securing the future: AI-driven cybersecurity in the age of autonomous IoT,” Frontiers in the Internet of Things, vol. 4, Art. no. 1658273, Sep. 2025. https://www.frontiersin.org/journals/the-internet-of-things/articles/10.3389/friot.2025.1658273/pdf
J. Sardana, R. Hariharan, and R. S. Pudi, “AI-driven secure and resilient healthcare data platforms with intelligent search and personalized communication,” in Proc. 2025 Tenth Int. Conf. Science Technology Engineering and Mathematics (ICONSTEM), Nov. 2025, pp. 1–6. https://ieeexplore.ieee.org/document/11374370
W. Yang and P. Chen, “eProbe: eBPF-Enhanced Accurate Container Status Probing in Cloud-Native Systems,” IEEE Transactions on Services Computing, vol. 18, no. 4, pp. 2047–2061, Jul.–Aug. 2025. https://ieeexplore.ieee.org/abstract/document/11023176/
H. Hanada and K. Ishibashi, “Service-Level Objective-Aware Load-Adaptive Timeout: Balancing Failure Rate and Latency in Microservices Communication,” IEEE Access, vol. 13, pp. 137881–137895, 2025. https://ieeexplore.ieee.org/iel8/6287639/6514899/11113294.pdf
H. D. Nguyen, T. Bicer, B. Nicolae, R. Kettimuthu, E. A. Huerta, and I. T. Foster, “Resilient execution of distributed X-ray image analysis workflows,” Frontiers in High Performance Computing, vol. 3, Art. no. 1550855, 2025. https://www.frontiersin.org/journals/high-performance-computing/articles/10.3389/fhpcp.2025.1550855/pdf
P. Oikonomou, K. Kolomvatsos, and C. Anagnostopoulos, “A proactive and uncertainty driven management mechanism for workflows of processing tasks,” Computing, vol. 107, Art. no. 112, 2025. https://link.springer.com/content/pdf/10.1007/s00607-025-01468-1.pdf
K. Priyanka, S. Priyadharshini, N. A. Vignesh, A. Hameed, K. R. Reddy, and C. Praveenkumar, “Self-Healing Data Pipelines: Reinforcement Learning for Real-Time Fault Detection and Autonomous Recovery,” in 2025 International Conference on Metaverse and Current Trends in Computing (ICMCTC), Subang Jaya, Malaysia, 2025. https://ieeexplore.ieee.org/abstract/document/11196544/
N. M. K. Koneru, “AI-driven optimization of CI/CD pipelines for multi-cloud infrastructure deployment,” in Proc. 2025 Tenth Int. Conf. Science Technology Engineering and Mathematics (ICONSTEM), Nov. 2025, pp. 1–6. https://ieeexplore.ieee.org/abstract/document/11374441/
L. Wang et al., “A survey on large language model based autonomous agents,” Frontiers of Computer Science, vol. 18, no. 6, Art. no. 186345, 2024. https://link.springer.com/content/pdf/10.1007/s11704-024-40231-1.pdf
M. Suzgun, T. Gur, F. Bianchi, D. E. Ho, T. Icard, D. Jurafsky, and J. Zou, “Language models cannot reliably distinguish belief from knowledge and fact,” Nature Machine Intelligence, vol. 7, pp. 1780–1790, 2025. https://www.nature.com/articles/s42256-025-01113-8
Y. Chang et al., “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology, vol. 15, no. 3, pp. 1–45, 2024. https://dl.acm.org/doi/pdf/10.1145/3641289
J. Shen, H. Zhang, Y. Xiang, X. Shi, X. Li, Y. Shen, Z. Zhang, Y. Wu, X. Yin, J. Wang, M. Xu, Y. Li, J. Yin, J. Song, Z. Li, and R. Nie, “Network-centric distributed tracing with DeepFlow: Troubleshooting your microservices in zero code,” in Proc. ACM SIGCOMM 2023 Conf., New York, NY, USA, 2023, pp. 420–437. https://dl.acm.org/doi/pdf/10.1145/3603269.3604823
L. J. Kohl et al., “Automated structural testing of LLM-based agents: Methods, framework, and case studies,” in Proc. 2025 IEEE Int. Conf. Big Data (BigData), Macau, China, 2025, pp. 1847–1856, doi: 10.1109/BigData66926.2025.11401679.
J. Saad-Falcon, O. Khattab, C. Potts, and M. Zaharia, “ARES: An automated evaluation framework for retrieval-augmented generation systems,” in Proc. 2024 Conf. North American Chapter Assoc. Comput. Linguistics: Human Language Technologies (NAACL-HLT), Mexico City, Mexico, 2024, pp. 338–354, doi: 10.18653/v1/2024.naacl-long.20.
F. Periti, H. Dubossarsky, and N. Tahmasebi, “(Chat) GPT v BERT dawn of justice for semantic change detection,” in Findings of the Association for Computational Linguistics: EACL 2024, Mar. 2024, pp. 420–436. https://aclanthology.org/2024.findings-eacl.29.pdf
A. Kaplunovich, “Advancing LLM agents for code generation: Observability, orchestration, reliable performance,” in Proc. 3rd Int. Conf. Intelligent Computing, Communication, Networking and Services (ICCNS), Varna, Bulgaria, 2025. https://ieeexplore.ieee.org/abstract/document/11428688/
M.-A. Storey, R. Hoda, A. M. P. Milani, and M. T. Baldassarre, “Guiding principles for mixed methods research in software engineering,” Empirical Software Engineering, vol. 30, no. 5, Art. no. 138, 2025. https://link.springer.com/content/pdf/10.1007/s10664-025-10629-x.pdf
A. S. Kang et al., “Enhancing biocompatibility: Heat-treated Al2O3- and ZrO2-reinforced hydroxyapatite coatings on metallic implants,” The International Journal of Advanced Manufacturing Technology, vol. 141, no. 12, pp. 6195–6208, 2025. https://www.researchgate.net/profile/Amishkumar-Patel/publication/398003329_Enhancing_biocompatibility_heat-treated_Al2O3-_and_ZrO2-_reinforced_hydroxyapatite_coatings_on_metallic_implants/links/69e7b402ef6ff80fdc762f05/Enhancing-biocompatibility-heat-treated-Al2O3-and-ZrO2-reinforced-hydroxyapatite-coatings-on-metallic-implants.pdf
O. Dhangar, “AI agentic systems with microservices architecture: Bridging software engineering lifecycle gaps in large-scale applications,” in Proc. 2025 IEEE Int. Conf. Recent Advances in Computing and Systems (REACS), 2025. https://ieeexplore.ieee.org/abstract/document/11413456/
C. Lohse, A. Selk, A. Ba, J. Wahl, and M. Ruffini, “Licence to Scale: A microservice simulation environment for benchmarking agentic AI,” in Workshop on Scaling Environments for Agents, Dec. 2025. https://openreview.net/pdf?id=03xThlPUxU
S. V. Vuddanti, A. Shah, S. K. Chittiprolu, T. Song, S. Dev, K. Zhu, and M. Chaudhary, “PALADIN: Self-correcting language model agents to cure tool-failure cases,” arXiv preprint arXiv:2509.25238, 2025. https://arxiv.org/pdf/2509.25238
B. Li, Y. Jiang, V. Gadepally, and D. Tiwari, “LLM inference serving: Survey of recent advances and opportunities,” in Proc. 2024 IEEE High Performance Extreme Computing Conf. (HPEC), Wakefield, MA, USA, 2024. https://arxiv.org/pdf/2407.12391
L. Zhu, Q. Zhuang, H. Jiang, H. Liang, X. Gao, and W. Wang, “Reliability-aware failure recovery for cloud computing based automatic train supervision systems in urban rail transit using deep reinforcement learning,” Journal of Cloud Computing, vol. 12, Art. no. 147, 2023. https://link.springer.com/content/pdf/10.1186/s13677-023-00502-x.pdf
B. Rombaut, S. Masoumzadeh, K. Vasilevski, D. Lin, and A. E. Hassan, “Watson: A cognitive observability framework for the reasoning of LLM-powered agents,” in Proc. 40th IEEE/ACM Int. Conf. Automated Software Engineering (ASE), Nov. 2025, pp. 739–751. https://arxiv.org/pdf/2411.03455
Q. Lu et al., “AgentArcEval: An architecture evaluation method for foundation model based agents,” Journal of Systems and Software, Art. no. 112656, 2025. https://arxiv.org/pdf/2510.21031
V. Nagaraj, S. R. Gundla, and P. R. Vennamaneni, “AI-driven secure and cloud-native microservice architectures for high-performance financial and hardware accelerator systems,” in Proc. 2025 Tenth Int. Conf. Science Technology Engineering and Mathematics (ICONSTEM), Nov. 2025, pp. 1–7. https://ieeexplore.ieee.org/document/11374740
T. Ou, W. Guo, A. Gandhi, G. Neubig, and X. Yue, “Agentdiagnose: An open toolkit for diagnosing LLM agent trajectories,” in Proc. 2025 Conf. Empirical Methods in Natural Language Processing: System Demonstrations, Nov. 2025, pp. 207–215. https://aclanthology.org/2025.emnlp-demos.15.pdf
C. Bonthu, G. Malik, and M. R. Dhanagari, “AI-driven secure and scalable data architecture with multi-domain MDM and cybersecurity for real-time fintech decision-making,” in Proc. 2025 Tenth Int. Conf. Science Technology Engineering and Mathematics (ICONSTEM), Nov. 2025, pp. 1–7. https://ieeexplore.ieee.org/document/11374416
P. Roshan, T. Prislan, M. Biswas, and A. Lee, “Semantic checkpointing for stateless LLM agents in multi-tenant enterprise systems,” 2025. https://www.researchgate.net/profile/Prakash-Roshan/publication/399433967_Semantic_Checkpointing_for_Stateless_LLM_Agents_Semantic_Checkpointing_for_Stateless_LLM_Agents_in_Multi-Tenant_Enterprise_Systems/links/695ac56b0c98040d482735c2/Semantic-Checkpointing-for-Stateless-LLM-Agents-Semantic-Checkpointing-for-Stateless-LLM-Agents-in-Multi-Tenant-Enterprise-Systems.pdf
Z. Wu, Q. Zeng, Z. Zhang, Z. Tan, C. Shen, and M. Jiang, “Large language models can self-correct with key condition verification,” in Proc. 2024 Conf. Empirical Methods in Natural Language Processing (EMNLP), Miami, FL, USA, 2024, pp. 12846–12867. https://aclanthology.org/2024.emnlp-main.714.pdf
S. Zhang, M. Yin, J. Zhang, J. Liu, Z. Han, J. Zhang, B. Li, C. Wang, H. Wang, Y. Chen, and Q. Wu, “Which agent causes task failures and when? On automated failure attribution of LLM multi-agent systems,” in Proc. 42nd Int. Conf. Mach. Learn. (ICML), vol. 267, 2025, pp. 76583–76599. https://arxiv.org/pdf/2505.00212?
V. Gundu, “Methodological foundations of AI observability for enterprise LLM applications,” International Journal of Engineering and Computer Science, vol. 14, no. 10, pp. 27742–27747, 2025.
S. Rangu, K. S. Chadha, and S. Samala, “AI-powered cloud architecture, agile analytics, and enterprise DevOps-enabled framework for healthcare and insurance ecosystem optimisation,” in Proc. 2025 Tenth Int. Conf. Science Technology Engineering and Mathematics (ICONSTEM), Nov. 2025, pp. 1–7. [https://ieeexplore.ieee.org/document/11374414
I. Bouzenia and M. Pradel, “Understanding software engineering agents: A study of thought-action-result trajectories,” in Proc. 40th IEEE/ACM Int. Conf. Automated Software Engineering (ASE), Seoul, Republic of Korea, 2025. https://arxiv.org/pdf/2506.18824
P. K. R. Prakashkumar, “Analytical study of Vertex–Oracle cloud integration in enterprise ERP systems for automated tax compliance and efficient financial closings,” International Journal of Accounting and Management Sciences, vol. 4, no. 4, 2025. https://scipubhouse.com/international-journal-of-accounting-and-management-sciences-ijams/content/ijams-vol-4-no-4-october-2025/analytical-study-of-vertex-oracle-cloud-integration-in-enterprise-erp-systems-for-automated-tax-compliance-and-efficient-financial-closings/
E reference: [1] H. Lee, S. Oh, J. Kim, J. Shin, and J. Tack, “ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification,” in Proc. 42nd Int. Conf. Machine Learning (ICML), vol. 267, 2025. https://arxiv.org/pdf/2502.14565
J. Sun, S. Y. Min, Y. Chang, and Y. Bisk, “Tools fail: Detecting silent errors in faulty tools,” in Proc. 2024 Conf. Empirical Methods in Natural Language Processing (EMNLP), Miami, FL, USA, 2024, pp. 14272–14289. https://aclanthology.org/2024.emnlp-main.790.pdf
N. AI, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. Gaithersburg, MD, USA: NIST Trustworthy and Responsible AI, 2024. https://pathologyinnovationcc.org/s/R8b_1722118549750.pdf
S. Zhang et al., “No more data silos: Unified microservice failure diagnosis with temporal knowledge graph,” IEEE Trans. Serv. Comput., vol. 17, no. 6, pp. 4013–4026, Nov.–Dec. 2024. https://nkcs.iops.ai/wp-content/uploads/2025/09/No_More_Data_Silos_Unified_Microservice_Failure_Diagnosis_With_Temporal_Knowledge_Graph.pdf
M. W. Ahmad, M. U. Akram, M. M. Mohsan, K. Saghar, R. Ahmad, and W. H. Butt, “Transformer-based sensor failure prediction and classification framework for UAVs,” Expert Systems with Applications, vol. 248, Art. no. 123415, 2024. https://www.sciencedirect.com/science/article/pii/S095741742400280X
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
All papers should be submitted electronically. All submitted manuscripts must be original work that is not under submission at another journal or under consideration for publication in another form, such as a monograph or chapter of a book. Authors of submitted papers are obligated not to submit their paper for publication elsewhere until an editorial decision is rendered on their submission. Further, authors of accepted papers are prohibited from publishing the results in other publications that appear before the paper is published in the Journal unless they receive approval for doing so from the Editor-In-Chief.
IJISAE open access articles are licensed under a Creative Commons Attribution-ShareAlike 4.0 International License. This license lets the audience to give appropriate credit, provide a link to the license, and indicate if changes were made and if they remix, transform, or build upon the material, they must distribute contributions under the same license as the original.


