ML Platform Engineer
Own model deployment tooling so research work reaches Plaid's financial platform without bespoke glue each time.
1638 roles across 160 companies — filter by region or category.
Showing 50 of 104 matching positions (51–100)
Clear filtersOwn model deployment tooling so research work reaches Plaid's financial platform without bespoke glue each time.
Build the internal platform that lets product teams ship to PlanetScale's developer platform safely and often.
Run the training and serving infrastructure behind PlanetScale's developer platform, including GPU scheduling, artefacts and rollouts.
Own Kubernetes, deployment tooling and paved-road abstractions underneath Postman's developer platform.
Build the pipelines, feature stores and monitoring that keep models behind Railway's developer platform reproducible and observable.
Own Kubernetes, deployment tooling and paved-road abstractions underneath Ramp's financial platform.
Own infrastructure-as-code, networking and capacity planning for Remote's employment platform.
Run the training and serving infrastructure behind Replit's developer platform, including GPU scheduling, artefacts and rollouts.
Own infrastructure-as-code, networking and capacity planning for Replit's developer platform.
Own infrastructure-as-code, networking and capacity planning for Rippling's employment platform.
Design and run the multi-region cloud footprint behind Scale AI's AI products, with cost and resilience both in view.
Own Kubernetes, deployment tooling and paved-road abstractions underneath Sentry's developer platform.
Own infrastructure-as-code, networking and capacity planning for Sierra's AI products.
Build the pipelines, feature stores and monitoring that keep models behind Snowflake's data platform reproducible and observable.
Own infrastructure-as-code, networking and capacity planning for Temporal's developer platform.
Build the pipelines, feature stores and monitoring that keep models behind Vercel's developer platform reproducible and observable.
Own infrastructure-as-code, networking and capacity planning for Weights & Biases's developer platform.
Build the internal platform that lets product teams ship to Wikimedia Foundation's programmes and platforms safely and often.
Run the training and serving infrastructure behind Zapier's product suite, including GPU scheduling, artefacts and rollouts.
Own Kubernetes, deployment tooling and paved-road abstractions underneath Zapier's product suite.
Own availability, latency and incident response for Atoms's robotics products, and remove the causes rather than the symptoms.
Automate the delivery pipeline behind Cal.com's open-source projects and keep deployments boring.
Set service level objectives and build the observability that makes ClickHouse's data platform debuggable under pressure.
Own availability, latency and incident response for Close's product suite, and remove the causes rather than the symptoms.
Set service level objectives and build the observability that makes CrowdStrike's security platform debuggable under pressure.
Own CI/CD, environments and release automation for Figma's design platform.
Own availability, latency and incident response for GitLab's developer platform, and remove the causes rather than the symptoms.
Set service level objectives and build the observability that makes Grammarly's product suite debuggable under pressure.
Own availability, latency and incident response for Gusto's employment platform, and remove the causes rather than the symptoms.
Automate the delivery pipeline behind Instacart's commerce platform and keep deployments boring.
Own CI/CD, environments and release automation for LangChain's developer platform.
Automate the delivery pipeline behind Linear's developer platform and keep deployments boring.
Own CI/CD, environments and release automation for Notion's product suite.
Own availability, latency and incident response for Palo Alto Networks's security platform, and remove the causes rather than the symptoms.
Own CI/CD, environments and release automation for PlanetScale's developer platform.
Own availability, latency and incident response for Postman's developer platform, and remove the causes rather than the symptoms.
Own availability, latency and incident response for Remote's employment platform, and remove the causes rather than the symptoms.
Set service level objectives and build the observability that makes Scale AI's AI products debuggable under pressure.
Own availability, latency and incident response for Snowflake's data platform, and remove the causes rather than the symptoms.
Keep a large multi-tenant DevSecOps platform reliable, working across Kubernetes, observability tooling and incident response.
Keep large-scale, multi-tenant observability infrastructure healthy across Kubernetes, Go and the company's own open-source stack.
Maintain the availability and performance of a large-scale observability platform running across multiple cloud regions.
Scale the serverless deployment infrastructure that builds and serves frontend applications globally.
Build and operate the Atlas managed database platform serving hundreds of thousands of clusters worldwide.
Design cloud-native security architectures for customers deploying workloads across AWS, Azure and GCP.
Keep Wikipedia and sister projects available worldwide by managing the infrastructure behind free knowledge.
Operate the cloud infrastructure platform serving droplets, managed databases and Kubernetes clusters at scale.
Operate the anycast network and edge infrastructure that routes traffic for a large share of the internet.
Scale the deployment infrastructure that builds, runs and auto-scales web services and background workers.
Lead the platform team scaling infrastructure for the incident management product across cloud regions.
No jobs match your search and filters. Try a different keyword or clear the filters.