Search
Browse By Day
Browse By Time
Browse By Person
Browse By Policy Area
Browse By Session Type
Browse By Keyword
Browse Artificial Intelligence Presentations
Program Calendar
Sign In
Search Tips
Policymakers who want to use evidence to inform their decisions face a complex task. They must assess whether a body of evidence is sufficiently credible, strong, and relevant to justify action at scale. In practice, this means answering a set of difficult questions: How confident should we be in research findings? Which programs have been shown to work? And are these programs applicable—and feasible—in a given policy context?
The tools available to support policymakers in answering these questions are limited. Evidence is often fragmented across sources, inconsistently reported, and difficult to compare. Even institutions designed to synthesize research into usable guidance—including evidence clearinghouses (e.g., CLEAR; CrimeSolutions; Economic Mobility Catalog; Social Programs that Work) and evidence quality ratings (e.g. GRADE Working Group, 2013; Ruggeri et al., 2020)—reach different conclusions about what counts as “good” or “strong” evidence. Even more importantly, they typically ignore contextual information like costs, political context, or implementation considerations, despite their importance for policymakers’ decision-making.
We build on existing frameworks to propose a three-step approach for identifying evidence that both “works” and is adoptable. In the first step, we distinguish credible evidence from research that cannot reliably evaluate a program’s effectiveness, relying on four research design criteria: evaluation methods, identifying assumptions, statistical power, and external validity. In the second step, we aggregate evidence across studies to determine whether a program is effective on the program-by-outcome level: we require at least two credible studies and more than 50% of those studies to agree on program’s effect. In the final step, we determine whether programs that “work” could also be adopted in practice. Because adoptability is highly context-specific, we propose availability of contextual and feasibility information – ranging from cost estimates to implementation considerations – as the single criterion.
We then use a large language model (LLM) to apply this approach to a corpus of policy research – nearly 16,000 research documents scraped from 27 evidence clearinghouses and policy research organizations – to assess not only the quality and strength of the evidence, but also whether adoptability information is reported. We find that around 14% of documents report credible studies and we identify 92 “effective” policy programs in our corpus. We also find that the availability of adoptability information varies widely across programs and that policymakers would be required to pull and synthesize contextual information from many documents to make adoption decisions.
Finally, we explore the extent to which the results of our landscape analysis are sensitive to the specific criteria we proposed and show that which programs are considered evidence-based varies widely. This also suggests that existing clearinghouses and classifications may arrive at different recommendations and complicate policymakers’ task of finding promising policy programs. Our three-step approach to identifying evidence on programs that work and are adoptable provides a new unifying framework for evidence synthesis. We also demonstrate that our criteria can be implemented at relatively low cost and at scale using LLMs, opening new avenues for researchers and policymakers to develop policy research databases that meet policymakers’ needs.