Senior Site Reliability Engineer
Own availability, latency and incident response for Close's product suite, and remove the causes rather than the symptoms.
914 roles across 160 companies — filter by region or category.
Showing 22 of 72 matching positions (51–72)
Clear filtersOwn availability, latency and incident response for Close's product suite, and remove the causes rather than the symptoms.
Set service level objectives and build the observability that makes CrowdStrike's security platform debuggable under pressure.
Own availability, latency and incident response for GitLab's developer platform, and remove the causes rather than the symptoms.
Set service level objectives and build the observability that makes Grammarly's product suite debuggable under pressure.
Automate the delivery pipeline behind Instacart's commerce platform and keep deployments boring.
Own CI/CD, environments and release automation for LangChain's developer platform.
Automate the delivery pipeline behind Linear's developer platform and keep deployments boring.
Own CI/CD, environments and release automation for PlanetScale's developer platform.
Own availability, latency and incident response for Remote's employment platform, and remove the causes rather than the symptoms.
Own availability, latency and incident response for Snowflake's data platform, and remove the causes rather than the symptoms.
Keep a large multi-tenant DevSecOps platform reliable, working across Kubernetes, observability tooling and incident response.
Keep large-scale, multi-tenant observability infrastructure healthy across Kubernetes, Go and the company's own open-source stack.
Scale the serverless deployment infrastructure that builds and serves frontend applications globally.
Build and operate the Atlas managed database platform serving hundreds of thousands of clusters worldwide.
Keep Wikipedia and sister projects available worldwide by managing the infrastructure behind free knowledge.
Operate the cloud infrastructure platform serving droplets, managed databases and Kubernetes clusters at scale.
Scale the deployment infrastructure that builds, runs and auto-scales web services and background workers.
Lead the platform team scaling infrastructure for the incident management product across cloud regions.
Build the deployment tooling and orchestration layer that lets developers ship apps globally in seconds.
Keep PostHog Cloud running reliably at scale across ClickHouse, Kafka and Kubernetes infrastructure.
Design the cloud infrastructure and platform tooling that keeps financial services reliable and compliant.
Operate and scale the managed database platform across multiple cloud regions with high availability targets.
No jobs match your search and filters. Try a different keyword or clear the filters.