Data Engineer, Evaluation
Build pipelines that collect, clean and version the evaluation datasets used to measure model quality before release.
538 roles across 160 companies — filter by region or category.
Showing 50 of 75 matching positions (1–50)
Clear filtersBuild pipelines that collect, clean and version the evaluation datasets used to measure model quality before release.
Own the indexing and retrieval pipelines that keep the answer engine fresh across the open web.
Build model training, serving and governance features that let enterprises run their own AI workloads on their own data.
Design and run evaluations that measure model capability and safety for enterprise and government customers.
Design red-teaming and evaluation methodology for the models behind Anthropic's AI products, and turn results into mitigations.
Design and run large-scale training experiments, then turn the findings into production model releases.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Booking.com's travel platform.
Optimise inference throughput and cost for the models running behind Booking.com's travel platform.
Train, evaluate and serve models that power Canva's design platform at production scale.
Optimise inference throughput and cost for the models running behind Chime's financial platform.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Clay's product suite.
Build the ingestion, storage and query layers that every team uses to reason about Clay's product suite.
Own the transformation layer and semantic definitions behind reporting on Cloudflare's developer platform.
Optimise inference throughput and cost for the models running behind Contentsquare's analytics platform.
Probe model behaviour for failure modes and build the evaluations that gate releases across Contentsquare's analytics platform.
Own the lakehouse and streaming infrastructure underneath Daraz Bangladesh's commerce platform, with an eye on cost and freshness.
Run experiments that push the model capabilities behind Databricks's data platform, and publish or ship what works.
Probe model behaviour for failure modes and build the evaluations that gate releases across Datadog's developer platform.
Train, evaluate and serve models that power Datadog's developer platform at production scale.
Own the lakehouse and streaming infrastructure underneath DoorDash's delivery network, with an eye on cost and freshness.
Own the lakehouse and streaming infrastructure underneath Figma's design platform, with an eye on cost and freshness.
Own the transformation layer and semantic definitions behind reporting on Freshworks's product suite.
Own the transformation layer and semantic definitions behind reporting on FusionAuth's security platform.
Investigate training, alignment and evaluation methods that make the models behind Glean's AI products more useful.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Groq's AI products.
Own the transformation layer and semantic definitions behind reporting on Gusto's employment platform.
Own the lakehouse and streaming infrastructure underneath Harvey's AI products, with an eye on cost and freshness.
Build the ingestion, storage and query layers that every team uses to reason about HubSpot's product suite.
Build the ingestion, storage and query layers that every team uses to reason about Klarna's financial platform.
Optimise inference throughput and cost for the models running behind Langfuse (ClickHouse)'s developer platform.
Own the lakehouse and streaming infrastructure underneath Lovable's AI products, with an eye on cost and freshness.
Design red-teaming and evaluation methodology for the models behind Mistral AI's AI products, and turn results into mitigations.
Own the modelling loop behind Nubank's financial platform — features, training pipelines, offline evaluation and online experiments.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Nubank's financial platform.
Own the transformation layer and semantic definitions behind reporting on Okta's security platform.
Own the lakehouse and streaming infrastructure underneath Palo Alto Networks's security platform, with an eye on cost and freshness.
Own the modelling loop behind Perplexity's AI products — features, training pipelines, offline evaluation and online experiments.
Probe model behaviour for failure modes and build the evaluations that gate releases across Perplexity's AI products.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Postman's developer platform.
Own the transformation layer and semantic definitions behind reporting on Retool's developer platform.
Own the transformation layer and semantic definitions behind reporting on Runway's AI products.
Own the lakehouse and streaming infrastructure underneath Runway's AI products, with an eye on cost and freshness.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Scale AI's AI products.
Build the ingestion, storage and query layers that every team uses to reason about ShopUp's marketplace.
Investigate training, alignment and evaluation methods that make the models behind Sierra's AI products more useful.
Build the ingestion, storage and query layers that every team uses to reason about Typeform's product suite.
Model user and system behaviour on Cloudflare's developer platform and turn the findings into decisions the team acts on.
Instrument, measure and interpret how people actually use Cloudflare's developer platform.
Build and operate the batch and streaming pipelines that move data across Datadog's developer platform.
Design experiments and causal analyses that decide what ships next across GoTo Group's delivery network.
No jobs match your search and filters. Try a different keyword or clear the filters.