With the Language Training and Testing Center (LTTC) and the NTNU Speech and Machine Intelligence Laboratory, developed an automatic GEPT speaking assessment system that combines pictures, questions, and speech.
System scope
Assesses spoken answers to task types such as picture description and question answering by combining the picture, the question, and the student's speech.
Integrates multi-aspect scoring from acoustic features, language use, image/question analysis, vision-language models, and LLMs.
Contributions and outcomes
Fine-tuned BLIP-2 and designed prompts for the T5 language model to improve multimodal speaking assessment; accuracy was 71% on familiar content and 68% on unseen content.
Published at O-COCOSDA 2024 and extended into the data augmentation method of the first-author SLaTE 2025 paper.
Collaborated with LTTC and the NTNU Speech and Machine Intelligence Laboratory; Advisor: Prof. Berlin Chen.