Individual Submission Summary
Share...

Direct link:

Assigning Proposals to Experts: Comparing Semantic Ranking and Local LLMs

Friday, November 6, 8:30 to 10:00am, Property: Boston Marriott Copley Place, Floor: 4th Floor, Room: Yardmouth

Abstract

Background
Submission volumes grow and assigning proposals to expert reviewers becomes an operational task with implications for quality, consistency, and equity. Automated reviewer assignment methods range from random assignments to text similarity to use of LLMs, but short proposal abstracts and heterogeneous reviewer expertise statements can introduce additional challenges. Recent studies show that matching quality between reviewers and proposals depends heavily on how expertise is represented and scored, and that different similarity algorithms can yield substantively different rankings even when applied to the same inputs (Stelmakh et al., 2023). In applied funding contexts, NLP- and LLM-based semantic matching systems are shown to closely replicate human-reviewed expert assignments than traditional approaches (Alvarez‑Garcia et al., 2026).
Purpose
This paper examines the LLM-supported matching approaches in comparison to keyword based and embedding-based semantic ranking for expert/reviewer-proposal assignment. We aim to compare approaches not only to matching performance but also to interpretability, operational usability, and transparency of results.
Data
Data includes 300 proposal abstracts (≈100 words each) and 52 reviewer information including reviewer-provided expertise descriptions and web search data on experts. The dataset reflects common real-world constraints: short proposal text, heterogeneous expertise language, and unstructured metadata.
Methods
We compare three approaches; a) keyword baselines that provide a transparent reference point, b) embedding based semantic ranking that encodes each proposal abstract and reviewer profile in a shared vector space and ranks reviewers per proposal by similarity score, c) LLM-supported approaches using a locally deployed LLM (Gemma 2) in two roles: first, LLM used for structured profiling that standardizes both proposals and reviewer profiles into comparable fields (e.g., topics, methods, applications), and second, LLM used for re-ranking that reorders top candidates retrieved via embeddings using a rubric. This comparison condition is motivated by evidence that explicit, structured profiling and rubric-based assessment can improve reviewer matching beyond implicit paper-to-paper similarity (Pan et al., 2024). There are multiple safeguards that are built into the pipeline: the LLM is open source and local and produces ranked recommendations with rationales.
Results
Preliminary results show that ranking stability and agreement patterns reflect prior findings that reviewer-similarity algorithms can be error-prone and sensitive to representation choices. Embedding-only ranking is faster and can be most practical for high-volume cycles (speed, consistency), while LLM-supported reranking is most informative with short text. 
Implications 
This study contributes a policy relevant comparison of proposal-expert matching options that organizations can select based on workload, risk tolerance, and interpretability needs. This work supports responsible adoption of AI by positioning LLMs as bounded decision support with human oversight.

References
Alvarez‑Garcia, E., Garcia‑Costa, D., De Waele, I., et al. (2026). Expert assignment system based on natural language processing for Marie Sklodowska-Curie actions. Scientific Reports, 16, 37115.
Pan, Y., Ning, Z., Wang, L., et al. (2024). Beyond paper-to-paper: Structured profiling and rubric scoring for paper–reviewer matching. arXiv.
Stelmakh, I., Wieting, J., Neubig, G., et al. (2023). A gold standard dataset for the reviewer assignment problem. arXiv.

Author