Infrastructure Engineer, Compute
Run the container and sandbox fleet that gives every user an instant, isolated development environment.
538 roles across 160 companies — filter by region or category.
Showing 29 of 29 matching positions (1–29)
Clear filtersRun the container and sandbox fleet that gives every user an instant, isolated development environment.
Design and run the multi-region cloud footprint behind Anthropic's AI products, with cost and resilience both in view.
Run the training and serving infrastructure behind Chime's financial platform, including GPU scheduling, artefacts and rollouts.
Own model deployment tooling so research work reaches Databricks's data platform without bespoke glue each time.
Build the internal platform that lets product teams ship to Datadog's developer platform safely and often.
Own infrastructure-as-code, networking and capacity planning for DoorDash's delivery network.
Own infrastructure-as-code, networking and capacity planning for Groq's AI products.
Own Kubernetes, deployment tooling and paved-road abstractions underneath Gusto's employment platform.
Build the internal platform that lets product teams ship to HubSpot's product suite safely and often.
Own Kubernetes, deployment tooling and paved-road abstractions underneath Notion's product suite.
Own infrastructure-as-code, networking and capacity planning for Okta's security platform.
Build the pipelines, feature stores and monitoring that keep models behind Okta's security platform reproducible and observable.
Build the internal platform that lets product teams ship to Palo Alto Networks's security platform safely and often.
Own Kubernetes, deployment tooling and paved-road abstractions underneath Perplexity's AI products.
Own Kubernetes, deployment tooling and paved-road abstractions underneath Postman's developer platform.
Own Kubernetes, deployment tooling and paved-road abstractions underneath Ramp's financial platform.
Run the training and serving infrastructure behind Replit's developer platform, including GPU scheduling, artefacts and rollouts.
Own infrastructure-as-code, networking and capacity planning for Replit's developer platform.
Design and run the multi-region cloud footprint behind Scale AI's AI products, with cost and resilience both in view.
Own infrastructure-as-code, networking and capacity planning for Sierra's AI products.
Own CI/CD, environments and release automation for Figma's design platform.
Own availability, latency and incident response for Gusto's employment platform, and remove the causes rather than the symptoms.
Own CI/CD, environments and release automation for Notion's product suite.
Own availability, latency and incident response for Palo Alto Networks's security platform, and remove the causes rather than the symptoms.
Own availability, latency and incident response for Postman's developer platform, and remove the causes rather than the symptoms.
Set service level objectives and build the observability that makes Scale AI's AI products debuggable under pressure.
Maintain the availability and performance of a large-scale observability platform running across multiple cloud regions.
Design cloud-native security architectures for customers deploying workloads across AWS, Azure and GCP.
Operate the anycast network and edge infrastructure that routes traffic for a large share of the internet.
No jobs match your search and filters. Try a different keyword or clear the filters.