Data Engineer, Evaluation
Build pipelines that collect, clean and version the evaluation datasets used to measure model quality before release.
1638 roles across 160 companies — filter by region or category.
Showing 50 of 229 matching positions (1–50)
Clear filtersBuild pipelines that collect, clean and version the evaluation datasets used to measure model quality before release.
Own the indexing and retrieval pipelines that keep the answer engine fresh across the open web.
Build model training, serving and governance features that let enterprises run their own AI workloads on their own data.
Design and run evaluations that measure model capability and safety for enterprise and government customers.
Train, evaluate and release open models, and maintain the libraries the community builds on.
Improve the fraud models scoring billions of payments a year without adding friction for honest buyers.
Model merchant lifecycles and turn the findings into product and pricing decisions.
Build the pipelines behind product analytics for services used by a very large share of the web.
Build semantic and hybrid search features on top of the Elastic Stack for enterprise customers.
Analyse acquisition and activation funnels and turn the results into decisions the team acts on.
Build the data platform behind transaction analytics for tens of millions of mobile wallet customers.
Train, evaluate and serve models that power Airbyte's open-source projects at production scale.
Build the ingestion, storage and query layers that every team uses to reason about Airbyte's open-source projects.
Own the lakehouse and streaming infrastructure underneath Amplitude's analytics platform, with an eye on cost and freshness.
Design red-teaming and evaluation methodology for the models behind Anthropic's AI products, and turn results into mitigations.
Design and run large-scale training experiments, then turn the findings into production model releases.
Own the transformation layer and semantic definitions behind reporting on Cursor (Anysphere)'s developer platform.
Build the ingestion, storage and query layers that every team uses to reason about Cursor (Anysphere)'s developer platform.
Run experiments that push the model capabilities behind Atoms's robotics products, and publish or ship what works.
Own the modelling loop behind Automattic's open-source projects — features, training pipelines, offline evaluation and online experiments.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Booking.com's travel platform.
Optimise inference throughput and cost for the models running behind Booking.com's travel platform.
Own the lakehouse and streaming infrastructure underneath BRAC's programmes and platforms, with an eye on cost and freshness.
Own the lakehouse and streaming infrastructure underneath BrowserStack's developer platform, with an eye on cost and freshness.
Own the modelling loop behind Buffer's product suite — features, training pipelines, offline evaluation and online experiments.
Train, evaluate and serve models that power Cal.com's open-source projects at production scale.
Build the ingestion, storage and query layers that every team uses to reason about Canonical's open-source projects.
Train, evaluate and serve models that power Canva's design platform at production scale.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Chaldal's commerce platform.
Optimise inference throughput and cost for the models running behind Chime's financial platform.
Investigate training, alignment and evaluation methods that make the models behind CircleCI's developer platform more useful.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Clay's product suite.
Build the ingestion, storage and query layers that every team uses to reason about Clay's product suite.
Design and run large-scale training experiments, then turn the findings into production model releases.
Optimise inference throughput and cost for the models running behind ClickHouse's data platform.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Close's product suite.
Own the transformation layer and semantic definitions behind reporting on Cloudflare's developer platform.
Own the modelling loop behind Cohere's AI products — features, training pipelines, offline evaluation and online experiments.
Design and run large-scale training experiments, then turn the findings into production model releases.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Coinbase's financial platform.
Own the lakehouse and streaming infrastructure underneath Coinbase's financial platform, with an eye on cost and freshness.
Design and run large-scale training experiments, then turn the findings into production model releases.
Optimise inference throughput and cost for the models running behind Contentsquare's analytics platform.
Probe model behaviour for failure modes and build the evaluations that gate releases across Contentsquare's analytics platform.
Own the lakehouse and streaming infrastructure underneath CrowdStrike's security platform, with an eye on cost and freshness.
Own the lakehouse and streaming infrastructure underneath Daraz Bangladesh's commerce platform, with an eye on cost and freshness.
Run experiments that push the model capabilities behind Databricks's data platform, and publish or ship what works.
Probe model behaviour for failure modes and build the evaluations that gate releases across Datadog's developer platform.
Train, evaluate and serve models that power Datadog's developer platform at production scale.
Optimise inference throughput and cost for the models running behind dbt Labs's data platform.
No jobs match your search and filters. Try a different keyword or clear the filters.