Individual Submission Summary
Share...

Direct link:

Uncertainty in Classifier Validation: Confidence Interval Estimation with Applications to LLMs and Nested Data

Saturday, November 7, 8:30 to 10:00am, Property: Boston Marriott Copley Place, Floor: 5th Floor, Room: New Hampshire

Abstract

Text classification – whether through supervised machine learning or large language models (LLMs) – allows researchers to extract measures from natural language data in large amounts and at limited expense. This capability is especially attractive for policy researchers who increasingly analyze legislative records, campaign speeches, public meetings, media reports, and social media postings (Brady, 2019; Grimmer & Stewart, 2013). However, resulting classifications are inevitably accompanied by error: some number of texts will be incorrectly classified into the category of interest while other true instances will be incorrectly classified as negatives. Thus, validation remains essential.   

The most common approach to validation is the presentation of performance metrics (like precision and recall) estimated within a sample of hand-classified texts (Grimmer et al., 2022; James et al., 2013). Although the magnitude of these metrics is essential for assessing the degree of trust that should be placed in resulting inferences, they are point estimates (estimated from a sample) with accompanying uncertainty. Yet, the reporting of confidence intervals surrounding performance metrics remains inconsistent in the social sciences (Anglin, 2024). Further, when intervals are reported, they are often computed using methods that are inaccurate for small sample sizes, high proportions, and/or nested data.

Using simulations and a case study reflecting scenarios common to text classification – including small to moderate effective sample sizes (e.g., due to a limited number of governing bodies under study), performance metrics approaching 1 (due to the high performance of LLMs), and non-independence (due to clustering of texts within authors), this study assesses best practices in the estimation and reporting of performance metric uncertainty. We evaluate coverage (the percent of samples whose intervals contain the true parameter) for six common 95% confidence intervals for proportion-based metrics: Wald, Wilson (1927), Agresti–Coull (1998), Clopper–Pearson (1934); bootstrap percentile (Tibshirani & Efron, 1993) and bootstrap BCa (Efron, 1987). We also propose and assess a novel pseudo-count regularized bootstrap, analogous to an Agresti-Coull interval but adapted for bootstrapping, with straightforward implementation and strong performance under simulation. For nested data, we simulate clustered populations (with performance metrics = 0.60 – 0.90; interclass correlation = 0.01–0.25) and compare unadjusted intervals to cluster-adjusted approaches using effective sample size, design degrees of freedom, and hierarchical bootstrapping (Dean & Pagano, 2015; Kish, 1965; Korn & Graubard, 1998). 

Results demonstrate that the commonly used Wald interval, percentile bootstrap, and BCa exhibit erratic and poor coverage at high proportions and small samples, while the Agresti-Coull, Clopper-Pearson, Wilson, and our proposed pseudo-count regularized bootstrap intervals reach 95% reliably across conditions. Under nested data, 95% confidence intervals assuming independence can obtain coverage as low as 45 percent. We illustrate these issues through a case study classifying segments of school board meeting minutes as public comment or not, a scenario where many texts are nested within a moderate number of authors: here, segments within meetings, within boards. We close with practical recommendations for reporting and sample size planning, with implications for the growing number of applications of LLM-based text analysis in policy research.

Author