Individual Submission Summary
Share...

Direct link:

Testing your way to impact: Iterative A/B testing for program optimization

Sat, March 28, 11:15am to 12:30pm, Hilton, Floor: Sixth Floor - Tower 3, Nob Hill 10

Proposal

Over the last six years, Youth Impact has conducted more than 75 A/B tests across our three education-related programs in Botswana, transforming how we optimize program effectiveness and cost-efficiency. This presentation shares practical lessons from embedding rapid experimentation into ongoing program operations, demonstrating how systematic A/B testing can simultaneously improve impact and reduce costs.

After successfully replicating our school-based HIV prevention program, Choices, through a randomized controlled trial in Botswana, we faced a familiar challenge: we had proven the program worked, but needed to optimize specific components for local context and scale. Traditional evaluation approaches would take years to provide answers we needed immediately for program expansion.

Beginning in 2018, we experimented with A/B testing as an alternative to lengthy RCTs. We started small, testing different messenger approaches and delivery formats within our health programming. By 2019, we had conducted 10 A/B tests and gained confidence to expand the methodology to our education programs: Teaching at the Right Level (TaRL) and ConnectEd phone-based tutoring.

Our ConnectEd program exemplifies the power of continuous testing. Through 12 successive A/B tests conducted every school term, we reduced program costs by 35 percent while improving learning outcomes—avoiding the typical "voltage drop" programs experience at scale. Each test built strategically on previous results.

Our testing trajectory illustrates how experiments cascade: when our initial RCT showed that phone calls plus text messages outperformed text messages alone, we used A/B testing to systematically optimize the intervention. For example, one test revealed that 40-minute calls twice monthly generated the same learning impact as 20-minute weekly calls—leading us to adopt the more cost-effective model program-wide. Another test showed that caregiver involvement significantly improved outcomes, prompting us to design subsequent experiments testing different strategies for encouraging parent engagement.

This iterative process shifted our evaluation focus from asking "Does the program work?" to "How can we make the program work better?" Tests examined both effectiveness enhancement and cost reduction, recognizing that sustainable impact requires both program improvement and financial viability.

Through extensive field experience, we developed a structured four-phase approach for organizations building A/B testing capacity, outlined in our Toolkit. The process begins with pilot implementation tweaks—practicing testing mechanics without randomization complexity—then progresses through establishing data flows, conducting first A/B tests, and maintaining ongoing experimentation cycles.
Critical to our success was identifying "golden indicators"—outcome measures that sit between easily-measured inputs (like teachers trained) and long-term impacts (like exam scores years later). Golden indicators respond to program changes within testing timeframes while remaining meaningful for program decisions. For our TaRL program, foundational learning assessments serve this role; for Choices, knowledge assessments provide rapid feedback while we monitor behavioral outcomes over longer periods.

Our experience demonstrates that A/B testing significantly increases program cost-effectiveness and scalability while building organizational capacity for evidence-based decision-making. However, success requires systematic approach, organizational commitment, and recognition that testing is most valuable when embedded within broader learning cultures rather than conducted as isolated activities.

Author