Search
Browse By Day
Browse By Time
Browse By Person
Browse By Policy Area
Browse By Session Type
Browse By Keyword
Browse Artificial Intelligence Presentations
Program Calendar
Sign In
Search Tips
Procurement is one of the most consequential administrative functions in government, yet it is often slow, document-intensive, and difficult for non-specialists to navigate. In city government, procurement work is governed by overlapping legal and procedural requirements, depends heavily on institutional knowledge, and frequently begins with staff who are not procurement experts. These features make procurement a useful setting for studying whether generative AI can improve public-sector work in a way that is both measurable and policy-relevant.
This paper evaluates BidBot, a context-specific procurement assistant developed by the City of Boston to answer questions about procurement rules and support the drafting of solicitation materials. Unlike general-purpose AI tools, BidBot is grounded in a curated body of local procurement guidance, templates, and related legal materials. The paper asks whether access to such a tool improves performance on common procurement tasks and whether any gains are concentrated in particular parts of the workflow.
We study these questions using a randomized field experiment with City of Boston employees engaged in procurement-related work. Participants were randomly assigned to complete a standardized procurement exercise either with access to BidBot or without it. The exercise included multiple-choice questions on procurement concepts, a specifications-writing task, a scope-of-work drafting task, and a post-exercise survey. We combine baseline survey data, task-level completion times, graded written outputs, and tool-usage logs. Written responses were evaluated using structured rubrics by procurement professionals and external policy practitioners.
We find that access to BidBot reduced total completion time, with the largest gains concentrated in the two writing tasks. By contrast, participants with access to the tool spent more time on the multiple-choice section, suggesting that the effects of AI vary across task types even within the same administrative domain. We also find meaningful improvements in the quality of scope-of-work documents, along with higher self-reported confidence in the quality of written outputs. Treated participants produced substantially longer written documents and reported more favorable beliefs about the usefulness and accuracy of generative AI for procurement work.
Taken together, the findings suggest that generative AI may be most useful in government when applied to bounded, rule-governed tasks that impose a high administrative burden and require substantial drafting effort. The results also point to an important distinction between tasks that benefit from AI-supported document production and tasks that continue to rely more heavily on human judgment. More broadly, the paper offers evidence on how local governments can evaluate and deploy domain-specific AI tools in ways that improve work quality while preserving institutional safeguards.