Search
Program Calendar
Browse By Day
Browse By Time
Browse By Person
Browse By Room
Browse By Unit
Browse By Session Type
Search Tips
Annual Meeting Registraion, Housing and Travel
Personal Schedule
Sign In
Advances in dialogue-based systems in areas such as ITS (Adamson et al., 2014; Graesser et al., 2001) make it possible to design conversations for assessment purposes. Conversation-based assessments (CBA) typically involve students interacting with one or more virtual characters. These CBAs provide students with multiple opportunities to demonstrate their knowledge/skills, are used to gather evidence of skills that may be difficult to obtain using traditional assessment approaches and could be used to elicit explanations about decisions that students make in other types of assessments (e.g., scenario and game-based assessments).
We have designed and implemented two CBA prototypes aimed at assessing science inquiry skills (Liu et al., 2016; Zapata-Rivera et al., 2014). The two isomorphic CBA prototypes that have been developed are the Volcano and Weather prototypes. These prototypes were designed to assess the following science inquiry skills: analyzing data and identifying patterns, conducting data collection, and making predictions based on data.
To develop these prototypes, conversation space diagrams (Zapata-Rivera, Jackson, Katz, 2015) were used to define the constructs for each conversation, provide the questions presented by the virtual characters, outline the expected categories of students’ responses to those questions (including NLP algorithms for classification – e.g., regular expressions, latent semantic analysis), and design subsequent prompts and follow-up responses for each category.
This presentation discusses issues related to the iterative development of the conversations used in these CBA prototypes, including: the complexity of regular expressions; how the complexity of regular expressions increases as new cases from evaluations are integrated into the system; limitations of the current approach; and ideas for future research and development in this area.
Data Sources
Data from the development of the Volcano and Weather prototype include different versions of the conversations obtained after conducting small- (path validation and cognitive lab studies) and large-scale studies (N=600, N=500, and N=300).
Results
Crowd sourced ratings from Amazon’s Mechanical Turk that passed calibration were utilized to improve the effectiveness of regular expressions. Regular expression modifications resulted in an overall reduction in the number of misclassified student responses in these datasets from 37% to 6% and 34% to 6% for the Volcano and Weather prototypes, respectively. Modifications to the regular expressions typically involved the inclusion of functional synonyms (i.e., words used synonymously in present context) and the identification of absent words/expressions. Ultimately, the regular expressions became more complex, but were also more effective in classifying student responses.
Discussion
A careful analysis of the number of errors fixed after making changes to the regular expressions can provide information on when it makes sense to make changes. Other strategies like adapting the responses of the virtual characters to cover a wide range of student utterances can be used to deal with some cases without increasing the complexity of the system. Also, the use of heuristics of guidelines when designing the conversation diagrams such as keeping the categories separate can help maintain the complexity of the conversations.
This work provides insights on how to design and improve automated conversations aimed at assessing science inquiry skills.
Diego Zapata-Rivera, Educational Testing Service
Blair Lehman, Educational Testing Service
Tanner Jackson, Educational Testing Service
Lei Liu, Educational Testing Service