Individual Submission Summary
Share...

Direct link:

Standardizing the Unstandardized: An NLP Pipeline for Private-School Transcript Research

Saturday, November 7, 10:15 to 11:45am, Property: Boston Marriott Copley Place, Floor: 5th Floor, Room: New Hampshire

Abstract

U.S. high-school course-taking research has largely drawn on state administrative data, where reporting requirements impose a common data frame across public schools. Independent and private schools sit outside that infrastructure. There is no centralized collection of private-school transcripts, and transcript files differ in layout even when obtainable. Linking records across private schools therefore requires rederiving a shared vocabulary for courses, subjects, and placement, work that has historically demanded intensive manual labor. This project treats standardization as the core methodological problem for bringing private-school course-taking data into education research at scale.

Transcript records have been used to study student academic behaviors, progression, Career and Technical Education (CTE) coursetaking, and curricular pathways (e.g., Hagedorn and Kress, 2008; Gottfried and Plasman, 2018; Malik, Feinberg, and Bruch, 2025), but largely in settings where a common data frame already exists, such as a single institution or administrative system. This project extends that work to a setting where no common frame exists, multiple independent schools with idiosyncratic transcripts, and asks what becomes possible when transcript PDFs are extracted and standardized into linkable records.

The paper addresses two questions. Can a generalizable pipeline, combined with a Natural Language Processing (NLP)-based standardization system, produce linkable transcript data across private institutions? And once such data exist, which cross-school analyses of course-taking, placement, and CTE participation become feasible?

I assemble transcripts from six schools spanning both text-based and image-based PDFs and build a three-stage Extract, Transform, and Load pipeline (PDF extraction, school-level standardization, and cross-school schema alignment) that produces unified student-level and course-level files. Building on this workflow, I am developing an agentic AI parser that uses structured prompting and output validation to interpret unfamiliar transcript layouts with limited human intervention. This component is still in progress.

The core methodological contribution is an NLP-based standardization system operating on three dimensions. First, course-name standardization: variants such as "US History," "U.S. Hist," and "American History" are mapped to canonical course identities through fuzzy matching and contextual embeddings. Second, course-subject classification: a supervised classifier built on DistilBERT, with SVM and logistic regression baselines, maps free-text course titles to academic tracks, reaching 98.5% accuracy across more than 212,000 records and 3,600 unique course names. Third, course-placement classification: the same framework is being extended to recognize honors, Advanced Placement, International Baccalaureate, and the 16 National Career Clusters that organize CTE coursetaking. Together, these layers turn a heterogeneous PDF archive into records that can be linked across institutions, a methodological innovation targeted at a sector underrepresented in education research precisely because such infrastructure has not existed.

The contribution of this paper is the extraction pipeline and the standardization system. With linkable transcript data in hand, downstream analyses such as course-sequence and pathway analysis, cross-school CTE participation, and placement patterns become tractable starting points for future work. The pipeline, agentic parser, and standardization system will be released as open-source tools, lowering the cost of working with private-school transcript microdata and bringing a long-absent sector into scope for education policy analysis.

Author