Paper Detail

Supporting Industrial Test-Failure Analysis with LLM-Based Systems: An Experience Report

Eric Jansson, Per Strandberg, Thomas Sörensen, Eduard Paul Enoiu, Wasif Afzal

arxiv Score 15.8

Published 2026-09-18 · First seen 2026-09-21

General AI

Abstract

This study examines tool-augmented Large Language Model (LLM) systems for supporting Root Cause Analysis (RCA) of nightly test failures at Westermo Network Technologies AB. Nightly test executions produce heterogeneous test data and logs that practitioners currently inspect manually across multiple sources. We implemented an RCA workflow in single-agent and orchestrated multi-agent configurations, both with access to test metadata and logs. An exploratory industrial case study used two real failure scenarios. Six practitioners evaluated the scenario reports through a survey and focus group, and operational measurements were collected from 120 repeated executions. The evaluation covered practitioner-perceived correctness, reasoning quality, fix realism, clarity, usefulness, and trust, as well as cost, duration, and consistency. Neither configuration showed a consistent practitioner-perceived quality advantage across the two scenarios. The single agent system generated reports faster and at lower cost, making it the more practical baseline in this context. The potential benefits of agent architectures require further evaluation in more complex scenarios.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
now
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@article{jansson2026supporting,
  title = {Supporting Industrial Test-Failure Analysis with LLM-Based Systems: An Experience Report},
  author = {Eric Jansson and Per Strandberg and Thomas Sörensen and Eduard Paul Enoiu and Wasif Afzal},
  year = {2026},
  abstract = {This study examines tool-augmented Large Language Model (LLM) systems for supporting Root Cause Analysis (RCA) of nightly test failures at Westermo Network Technologies AB. Nightly test executions produce heterogeneous test data and logs that practitioners currently inspect manually across multiple sources. We implemented an RCA workflow in single-agent and orchestrated multi-agent configurations, both with access to test metadata and logs. An exploratory industrial case study used two real fail},
  url = {https://arxiv.org/abs/2609.21843},
  keywords = {cs.SE},
  eprint = {2609.21843},
  archiveprefix = {arXiv},
}

Metadata

{}