The 17th IEEE International Conference on Software Testing, Verification and Validation (ICST) took place in Toronto, Canada, from 27 to 31 May 2024. The conference attracted researchers, engineers and industry practitioners who presented new findings in software testing and quality assurance. Among the events was a first‑ever workshop titled “Artificial Intelligence in Software Testing” (AIST), which focused on the use of large language models (LLMs) in testing activities. A panel discussion that followed the workshop addressed the reliability of LLMs and the expectations that can be placed on them.

According to the author of this article, the panel centered on two questions: how can research on LLMs be conducted reliably, and how much benefit can realistically be expected from these models. The author noted that, at the time, the scale of potential harms from LLMs was not yet well understood. The discussion highlighted that empirical studies of LLM effectiveness are difficult to generalise, and that results could be challenged by arguing that the wrong targets or applications were studied.

While the workshop was a key part of the conference for the author, the most impactful session was the keynote by Mike Hoye, titled “We build the world we measure.” Hoye criticised the software industry for its perceived disregard for customer well‑being and for the lack of accountability in software development. He called for accountability mechanisms similar to those used in other engineering disciplines, arguing that software is too integral to daily life to be treated as a low‑risk activity.

During the keynote, the author recalls a brief exchange with two Google engineers—one senior and one junior—who asserted that accountability is toxic and undermines software freedom. The engineers repeated their position without providing evidence beyond the keynote’s arguments. The author described the interaction as frustrating, noting that the engineers’ stance reflected a broader industry tendency to treat consent as optional, responsibility as a warranty limitation, and negligence as excusable in the name of progress.

The author connects these observations to the rise of LLM‑assisted development. The narrative suggests that LLMs enable any programmer to generate code that can be commercialised, potentially reducing the need for deep understanding and accountability. The author argues that this shift could make developers more expendable, as the tools that produce code become the focus of responsibility rather than the individual programmers. The author warns that, in a future where product designers generate code directly, no single person may feel accountable for failures, leading to a culture where software is expected to fail, even when critical to users.

The author concludes that companies that create and distribute software must be held accountable. The article notes that historically, software quality has depended on programmers striving for high standards, but that dysfunction has become common. The author stresses that users are currently suffering from negligent practices and that software quality is likely to degrade further without accountability.

The workshop was the first of its kind, and the author delivered a tutorial on a potential competition for AI‑assisted testing. The panel discussion, which evolved into a roundtable, focused on the reliability of LLMs and the realistic expectations for their application. The author, who has expressed skepticism toward LLMs, noted that the scale of potential harms was not yet well understood at the time.

ICST 2024 attracted researchers, engineers, and industry practitioners from around the globe. The conference program included over 50 peer‑reviewed papers and a range of keynote speeches, with the AIST workshop standing out as a new focus area for applying generative AI to software testing.

The author observes that the software industry has long treated consent as optional, responsibility as a warranty limitation, and negligence as excusable in the name of progress. This culture, the author argues, is reinforced by the rapid adoption of LLMs, which can produce code with little human oversight.

If product designers can generate code directly using LLMs, the author warns that accountability could become diffuse. In such a scenario, no single individual would feel responsible for failures, potentially normalizing software defects even in critical systems.

The ICST 2024 conference highlighted the growing debate over AI‑driven software accountability. The author’s reflections underscore the need for clear accountability mechanisms as LLMs become more integrated into development workflows. Without such measures, the industry risks a future where software quality deteriorates and users bear the consequences of unchecked automation.