•
BS, MS, or PhD in Computer Science or a related technical field, or equivalent practical experience
•
Hands-on experience operating production Kubernetes clusters — node lifecycle, upgrades, troubleshooting — plus GitOps and infrastructure-as-code experience
•
Experience with GPU or ML workload scheduling, queueing and priorities, fractional GPU sharing, autoscaling, or multi-tenant resource management
•
Self-driven with a strong sense of ownership: a quick learner who is eager to take responsibility and drive projects forward end to end
•
Experience with Ray or Kubeflow
•
Experience with lakehouse technologies such as Delta Lake or Apache Iceberg
•
Experience operating large-scale distributed data-processing and workflow systems, with hands-on depth in a system such as Apache Spark and working knowledge of Argo Workflows or an equivalent orchestrator