Senior Site Reliability Engineer
Own availability, latency and incident response for Close's product suite, and remove the causes rather than the symptoms.
914 roles across 160 companies — filter by region or category.
Showing 24 of 74 matching positions (51–74)
Clear filtersOwn availability, latency and incident response for Close's product suite, and remove the causes rather than the symptoms.
Set service level objectives and build the observability that makes CrowdStrike's security platform debuggable under pressure.
Own availability, latency and incident response for GitLab's developer platform, and remove the causes rather than the symptoms.
Set service level objectives and build the observability that makes Grammarly's product suite debuggable under pressure.
Own CI/CD, environments and release automation for LangChain's developer platform.
Automate the delivery pipeline behind Linear's developer platform and keep deployments boring.
Own CI/CD, environments and release automation for PlanetScale's developer platform.
Own availability, latency and incident response for Remote's employment platform, and remove the causes rather than the symptoms.
Own availability, latency and incident response for Snowflake's data platform, and remove the causes rather than the symptoms.
Own availability, latency and incident response for Tailscale's developer platform, and remove the causes rather than the symptoms.
Keep a large multi-tenant DevSecOps platform reliable, working across Kubernetes, observability tooling and incident response.
Deploy and troubleshoot Ubuntu, OpenStack and Kubernetes environments alongside enterprise customers.
Keep large-scale, multi-tenant observability infrastructure healthy across Kubernetes, Go and the company's own open-source stack.
Build and scale the cloud platform infrastructure supporting all Atlassian products across global regions.
Scale the serverless deployment infrastructure that builds and serves frontend applications globally.
Keep thousands of managed Postgres instances healthy across global regions with automated tooling.
Build and operate the Atlas managed database platform serving hundreds of thousands of clusters worldwide.
Scale the infrastructure supporting flash sales and peak traffic events for the global commerce platform.
Keep Wikipedia and sister projects available worldwide by managing the infrastructure behind free knowledge.
Operate the cloud infrastructure platform serving droplets, managed databases and Kubernetes clusters at scale.
Lead the platform team scaling infrastructure for the incident management product across cloud regions.
Build the deployment tooling and orchestration layer that lets developers ship apps globally in seconds.
Keep PostHog Cloud running reliably at scale across ClickHouse, Kafka and Kubernetes infrastructure.
Operate and scale the managed database platform across multiple cloud regions with high availability targets.
No jobs match your search and filters. Try a different keyword or clear the filters.