Background
- Started in 2025-10; I own the architecture, the SvelteKit frontend and backend, Azure Speech and Azure OpenAI integration, the PostgreSQL data flow, and deployment.
- Task types are short-passage reading, dialogue reading, long-passage reading, question answering, picture description, information response, and opinion expression; it offers Chinese and English interfaces, signed sessions, rate limiting, input sanitization, and audit logs.
- Flow: the browser records audio and uploads it as 16 kHz WAV, and scoring jobs enter BullMQ; a worker first scores pronunciation and fluency with Azure Speech, then scores content and generates feedback with Azure OpenAI, and finally aggregates the result.
Async scoring: surviving rate limits, restarts, and simultaneous exam starts
- Problem: A scoring job originally ran once and failed permanently on any Azure rate limit or timeout; a stress test also showed that a deployment restart re-enqueued pending jobs as duplicates.
- Approach: Moved to up to 4 attempts with hour-scale exponential backoff, deduplicated by using the answer ID as the job ID, and replaced an existing job only on a manual admin retry; also built a load-testing tool that simulates an entire exam.
- Result: A load test at the scale of 2000 sessions had 0 failures. The same round caught a problem where, with 100 students starting an exam at once, about 50% of requests failed; it was fixed by switching to READ COMMITTED with a unique constraint and retries.
Recording reliability: finding the real cause with audit data
- Problem: Some students' recordings arrived empty, but the frontend only logged errors to the console, leaving nothing to investigate.
- Approach: Reused the microphone stream obtained during the device check for the whole exam, negotiated the recording format each browser actually supports, and wrote every recording's format, size, peak volume, and error to the audit log; the device check also gained a live waveform and rejects recordings that are too quiet.
- Result: Audit data showed 165/291 WebKit webm recordings were only 5 bytes, disproving the original "cold start" hypothesis; switching WebKit to mp4 fixed it and identified 47 affected iPad answers from the exam.
Scoring fairness: keeping speech-recognition errors from becoming violations
- Problem: Speech recognition misheard read-aloud content and triggered the AI service's content filter; the old flow discarded the already-computed pronunciation score and could even flag the whole attempt as a violation, and 11 read-aloud answers were misjudged this way while live.
- Approach: Kept the speech score when feedback generation fails and separated the model's own filtered output from student violations; also rebalanced speech and content weights for teaching needs, with a backfill script that needs no new AI calls.
- Result: Speech-recognition errors no longer flag an attempt or lose its score, and scoring-rule changes can be applied directly to existing results.
