System scope
- Competition: 2026 Yunyong Zhisheng · 1111 Smart Job Search, where the retrieval system scored second.
- The fixed dataset contains 1,218,635 job postings and provides keyword search, filtering, sorting, and a job detail API.
- Tantivy BM25 is the production retrieval core; Qwen embedding, reranker, multi-view, and Graph retrieval are offline experiments that must pass a promotion gate before going live.
- The SvelteKit Web app and FastAPI API were deployed on one domain, with data in Aurora PostgreSQL.
Contributions and outcomes
- Built a one-command fail-closed pipeline from safe competition ZIP extraction and 39-column schema/taxonomy/SHA-256 validation through Aurora import, index building, and deployment.
- Designed an immutable runtime manifest and content-addressed S3 artifacts so dataset, index, model, and image versions are traceable and cannot be mixed.
- Built retrieval ablation and a promotion gate; a challenger can be enabled only when fixed-input NDCG@10 evidence is positive.
- Completed the production pipeline with GitHub OIDC, AWS CDK, ECS, CloudFront, and SageMaker, adding image scan, readiness, ranking, job detail, and Web UI public smoke checks.
Screenshot provenance: The original SvelteKit frontend runs locally, showing search, filtering, and result cards with three fixture jobs conforming to the production API contract; these are not actual search results from the production dataset.

