Owns the AI-collaboration scoring in the AI Skill Engine and the end-to-end Skill Challenge flow, from question design and the answering environment to scoring and reports.
Problem: When assessing how developers work with AI coding agents, scoring a whole conversation in one pass squeezed scores into a narrow band that could not separate experts from beginners; tuning the prompt improved known samples but not held-out ones.
Approach: Switched to turn-by-turn scoring, giving each turn its preceding turns as context, and built an independent validation benchmark—synthetic user personas, reproducible perturbations, and A/B comparisons—so that a regression is caught and rolled back.
Result: On real data, the gap between model and human scores falls within the gap between two human raters; the design trade-offs and validation are recorded as ADRs.
Skill Challenge: end to end, from question design to reports
Problem: When questions, answering environments, and grading criteria are maintained separately, they drift apart and results become hard to trace.
Approach: Designed a question-authoring process gated by staged human reviews and trial runs by an AI agent; decomposed each challenge's checkpoints into minimal, discriminating assertions and turned them into contracts the grading engine executes directly; bound submissions to the original files by hash so the grading evidence is traceable.
Result: Delivered the full flow from question design and answering environment to scoring and reports, with automatic submission intake and sync through Google Classroom.
Also built a local collection tool for AI coding agent transcripts that de-identifies data before export, validated on macOS, Linux, and Windows.