Maintains the lab's Linux GPU hosts and container environments, handling equipment operations with Ansible, SSH, and batch management tools; user environments cover Docker/Podman, Apptainer, and HPC/SLURM.
Independently built the lab's GPU host monitoring and account management system: a FastAPI and React admin interface, Prometheus and Alertmanager alerts pushed to Discord, IPMI power control, and account changes executed through Ansible.
Completed an account and GPU asset audit of 29 hosts, finding long-unused accounts and reconciling the equipment inventory.