RL Environment Engineer
Bespoke Labs
- Developed AI agent evaluation tasks for Nebula Aurora, a Kubernetes-based benchmarking platform that tests frontier AI models on real-world DevOps and SRE scenarios
- Wrote setup scripts that inject controlled failures into K3s clusters, simulating infrastructure incidents across CI/CD pipelines, service mesh configurations, and observability stacks
- Built Python graders with partial scoring logic to assess AI agent responses against actual system state, covering incident response, platform engineering, and cloud operations tasks
- Authored solution scripts as ground-truth references; each task had to be fully verifiable from observable system state alone