Background
- Started in 2026-03 with a team of 3 today; as founder and main maintainer I own the overall architecture, the judging system, and production releases.
- Supports eight languages—C, C++, Go, Java, JavaScript, Python, Rust, and TypeScript—and three judging modes: Standard, Checker, and Interactive; provides ICPC/IOI scoring, a live scoreboard, freeze, course assignment deadlines, and Dolos AST code-similarity detection.
- Flow: a submission is written to PostgreSQL and object storage, then a Temporal workflow dispatches it to a judge worker; each judging stage runs in a gVisor sandbox, and the verdict returns to the browser over SSE via Redis—if SSE drops, the frontend falls back to polling, so no result is lost.
Judge scheduling: replacing a custom coordinator with Temporal's native priorities
- Problem: During a rejudge of 789 submissions, every queued workflow kept polling a custom capacity coordinator; workers were busy replaying event history, the queue backed up by about 700 tasks for 16 minutes, and students' live submissions were stuck behind the rejudge.
- Approach: Switched to Temporal's native priorityKey and fairnessKey so exams outrank contests, which outrank practice, recovery, and rejudges, with at most one submission per student being judged at a time; capacity now comes directly from worker slots with a Kubernetes ResourceQuota as the hard cap, which let the custom coordinator be deleted entirely. Kueue and HPA/KEDA were evaluated but not worth it on a single-node cluster.
- Result: In a load test, 100 submissions were all AC within 243 seconds; after deleting the pods of Temporal's 4 component types one by one, 32 submissions were still all AC with no lost work.
Sandbox: measuring time and memory accurately
- Problem: Timing by the whole container's cgroup CPU charged the runner's own overhead to student code, so an empty program measured 30–60 ms; gVisor also lacks memory.peak, so a memory overrun could OOM the whole container.
- Approach: Wrote nojv-exec, about 170 lines of C that limits resources with rlimit and takes CPU time and peak RSS from wait4, with wall-clock time only as a watchdog; moved to one Pod per stage with separate prepare, run, and judge containers, keeping expected answers only in the judge container. Heavier isolation such as Firecracker/Kata was evaluated, but I kept gVisor and focused the effort on measurement itself.
- Result: An empty program now measures 0.98 ms, and the memory-limit smoke test went from about 1/3 failing to 6/6 passing. I later restricted IPC with seccomp and cleared scratch space before and after each test case, closing state leaks between test cases within a stage.
Release and reliability
- A tag triggers the build; images carry attestations and are written by digest to a deploy branch, Flux deploys them to Kubernetes, and an external status service verifies release, livez, and readyz.
- Sandbox cleanup now deletes with UID preconditions and durably retries anything it cannot confirm; after fixing a Kubernetes client that opened a new TLS connection on every call, cleanup latency p50 dropped from about 541 ms to 121 ms.
- Built reference solution validation that actually verifies problems, answers, and runtime environment changes before merging.
Screenshot provenance: Real product screens of the signed-in problem library and Monaco Editor.


