Search
Browse By Day
Browse By Time
Browse By Person
Browse By Policy Area
Browse By Session Type
Browse By Keyword
Program Calendar
Sign In
Search Tips
The accelerating use of Artificial Intelligence (AI) in knowledge work is replacing experts in analyzing or synthesizing organizational information. There is, however, a limited understanding of the impact that occurs when AI is introduced into a work process at scale. Moreover, analyses typically overlook the contextual differences between use cases, such as whether it is being used for financial analysis, customer service, or chatbots. The assessment of AI impacts needs to be disaggregated and evaluated for its outcomes in specific application contexts, tasks, and workplace interactions before broader impacts across organizations and occupations can be understood.
We conducted a case study of the United Nations Development Programme’s (UNDP) use of AI for quality assessment (QA) of project evaluation reports, which provides feedback to country offices for project improvement. To reduce operational costs from its annual project budget of $5 billion, its Independent Evaluation Office piloted a project to replace subject matter experts (SMEs) for QA with an AI tool based on a general-purpose large language model (LLM).
To determine how the AI tool performs compared to SMEs, and whether human oversight remains necessary, we analyzed over 750 AI-generated QA ratings and comments, along with data from an additional evaluation step reviewing AI outputs against prior SME comments and original report content. This data includes scores and descriptive feedback from human evaluators assessing AI comments using the following criteria: narrative structure, analytical depth, usage of terminology, consistency with AI-generated ratings, and factual inaccuracies.
Our study includes three core analyses. First, for a quantitative analysis, we aggregated scores across comment assessment criteria. Second, for a qualitative content analysis, we identified recurring issues flagged in the initial evaluation. Finally, we conducted a linguistic analysis comparing AI and SME comments using word frequency distributions, semantic diversity, sentiment analysis, sentence length, and punctuation use.
The AI tool performed well based on the quantitative analysis, but the content analysis showed discrepancies between AI ratings and corresponding comments. AI outputs also showed structural narrative issues, overstated report quality, and generated poorly grounded assessments. The linguistic analysis corroborated these findings. Although the AI comments mimicked the statistical distribution of SME comments, our most compelling findings show AI comments had significantly lower semantic diversity across a wide range of reports, indicating a lack of context-specific assessment. Furthermore, sentiment analysis revealed an overly positive bias that reflected the overstatement of report quality and divergence from generated ratings.
Rather than early-adopter hurdles, our findings suggest intrinsic constraints in general-purpose LLMs prevent the complete automation of this QA task, requiring a human-in-the-loop to identify when outputs contain inconsistencies and lack substantive depth. Rather than automating QA in isolation, we recommend using AI as an assistive tool to manage routine compliance tasks, allowing SMEs to focus on the substantive parts of the process and country offices to use it for pre-screening reports. Incorporating AI into knowledge work at scale operates as a broader transformation of the work process, with outcomes depending on the tradeoffs and evolutionary decisions about workflow and its purpose.