Applications are closed for this role. It was originally posted 2026-09-21. It’s no longer accepting applicants — see roles Tracksuit Limited is still hiring for →, or browse the live openings below.
About this role
- 🧱 Platform-wide impact: You set how reliability, observability and operational readiness work across the whole platform, and every team shipping on it feels the difference.
- 🤖 Agentic systems in production: You’ll set the guardrails, identity and blast-radius controls for agents operating against real systems and data, and work out what could go wrong before it does.
- 🛠 Golden paths: You’ll make the platform legible to agents as well as people, through golden paths, MCP servers, skills and documentation that both can actually use.
- 🚀 Startup pace with scaleup reach: Five years in, we’re working with over 1,000 brands across AU, NZ, USA and the UK, with plenty of scaling still ahead.
- Build and maintain the cloud infrastructure and paved paths teams ship on, so resilience, security and scalability come as defaults rather than as decisions each team makes on its own
- Give teams observability they can self-serve: monitoring, alerting, logging and tracing, plus automation for provisioning, deployments and the operational work nobody should be doing by hand
- Make incidents easier to handle: the tooling, runbooks and practice that let whoever is closest to the problem debug it quickly, and post-incident reviews that turn into actual changes
- Set the identity and least-privilege model for non-human actors, including agents, CI and MCP servers, so teams can give agents real access without real risk
- Document infrastructure designs and operational procedures in a form both people and agents can act on, including machine-readable runbooks
- Set the standard for what is safe to ship, and make it easy to meet through guardrails and checks in the pipeline rather than through gatekeeping
- Give teams visibility and controls over cloud and inference spend, so the cost of what they run is something they can see and act on
- Coach engineers in reliability practices, and work with Engineering and Product to balance feature delivery against platform needs
- You’ve run incident command on real production incidents, and you’ve owned a platform through a meaningful scaling step
- Deep infrastructure skills: Infrastructure as Code (Terraform, Terragrunt, CDK), containers (ECS), cloud platforms (AWS preferred, or GCP/Azure), CI/CD, and scripting in Python, TypeScript or Bash.
- Strong on observability: Datadog or similar, distributed tracing and structured logging, and a solid grasp of incident management methodology.
- Security-minded: Networking and cloud architecture fundamentals, plus an understanding of the security model for agentic systems, including credential handling, least privilege and prompt injection.
- Works with agents: You use coding agents in your own work (Claude Code or equivalent) and have a point of view on where they help and where they don’t.
- Product-oriented: You navigate technical complexity with business outcomes in mind, and can clearly communicate system health, risks and trade-offs to technical and non-technical people.
- Kind: This is a collaborative role, so we ultimately want a great, supportive team player.
- Bonus points if you’re passionate about marketing, design, and building exceptional user experiences.
Some of the tools we use: AWS, ECS, Terraform and Terragrunt, GitHub Actions, Datadog, Claude Code, Linear, Notion, Postgres, DynamoDB and Snowflake.
We practice transparent compensation at Tracksuit, which means we put a real focus on fair compensation for our people and their roles. We run comprehensive full-cycle comp reviews annually, as well as rolling promotions when you've levelled up. Your salary is supplemented with our best-in-class benefits package, as well as generous ESOP.
Tech stack
TerraformAWSGCPAzurePythonTypeScript
Similar roles at other companies
2
Senior Site Reliability Engineer
Engineering
$96K–$142K
B
Senior Site Reliability Engineer - FedRAMP
Engineering
$140K–$185K
M
Site Reliability Engineer (Mid-Level, Senior or Staff), Infrastructure Security
Engineering
$127K–$249K
N
Senior Site Reliability Engineer
Engineering
$60K–$85K
See all Engineering roles →