AI Evaluation Tools Must Be Transparent, Warns UW Philosopher
Lee’s central concern is that many AI‑based evaluation tools are being applied beyond the scope for which they were created. She refers to this misuse as “inference by false ascent,” a term coined by Nancy Cartwright and colleagues. The fallacy occurs when a metric built for a concrete purpose is elevated to a more abstract concept and then repurposed in ways that the original designers did not intend. Lee cites the journal impact factor as a historical example: a tool created to help librarians decide which journals to purchase was later used as a proxy for scientific impact and, more recently, for individual researcher quality.
The problem, Lee argues, is amplified in AI because the field is still developing norms for what counts as a reliable measure of scientific quality. Technical demands often prioritize large training datasets over epistemic fit, and there is widespread disagreement about how to conceptualize and measure normative aspects of science. These factors make AI evaluation tools especially vulnerable to false ascent.
To counter this, Lee proposes a new scientific norm she calls “Critically Engaged Pragmatism.” The norm would require scientific communities to scrutinize the purpose and purpose‑specific reliability of AI evaluation tools. It would also demand that tool creators provide full disclosure of design choices, training data, and benchmarking procedures so that independent reviewers can assess reliability, error liability, and bias.
Lee points to existing transparency initiatives as a model. The Consolidated Standards of Reporting Trials (CONSORT) and the Transparency and Openness Promotion (TOP) Guidelines are examples of living documents that evolve as new evidence emerges. She argues that similar standards for AI evaluation tools should be updated more frequently, perhaps in a living‑document format, to keep pace with emerging forms of bias, error, and gaming.
The article also notes that open‑weight large‑language models (LLMs) offer a potential path forward. Because their weights and code are publicly available, researchers can audit and customize them for specific scientific domains. Lee suggests that the scientific community should invest in developing and evaluating LLM‑based agents that are transparent and controllable, rather than relying on proprietary commercial systems that are difficult to scrutinize.
While Lee acknowledges that the full implementation of Critically Engaged Pragmatism faces structural obstacles—such as the divide between academia and industry, and limited access to technical expertise—she emphasizes that the norm provides a clear direction for future policy and practice.
The article concludes that AI evaluation tools are not objective arbiters of scientific credibility. Instead, they are the subject of critical discourse that ultimately underpins the credibility of scientific communities. By demanding transparency and purpose‑specific reliability, the scientific community can better guard against the unintended consequences of AI tools.
The discussion comes at a time when the use of AI in research is expanding rapidly. Several journals and conferences are already exploring requirements for submitting code and log traces alongside AI‑generated manuscripts, a move that aligns with Lee’s call for transparency. The broader AI ecosystem is also seeing increased attention to open‑source and open‑weight models, which may help meet the standards she outlines.
In short, Lee’s article urges that the scientific community adopt a disciplined, transparent approach to AI evaluation tools, ensuring that these tools serve their intended purpose without overstepping into domains where they have not been properly validated.