Staff Engineer, Site Reliability
Staff Site Reliability Engineer responsible for leading platform resilience and observability work. Role sits in the Site Reliability Engineering team within Engineering to scale infrastructure and enable observability and self-service tooling.
What you'll do
- Identify infrastructure improvements for performance, cost, and maintainability
- Design and lead an observability function for metrics, logs, transactions
- Drive resilience and scaling strategies across engineering teams
- Build monitoring, alerting, and self-service tooling for engineers
- Participate in on-call rota and mentor junior engineers
What they're looking for
- 7+ years in software or operations roles
- 5+ years cloud engineering experience, 2+ years with AWS
- Experience with Kubernetes and Docker microservice deployments
- Experience designing and implementing observability stacks and SLO/SLI frameworks
- Experience with IaC, automation tooling, and CI/CD systems
Skills
Summary written by StartupJobs from the company's listing. The full listing is on the company's website.
This job comes from the career page of LearnUpon. We link straight to the source.