Data Engineer, Evaluation
Build pipelines that collect, clean and version the evaluation datasets used to measure model quality before release.
1638 roles across 160 companies — filter by region or category.
Showing 50 of 150 matching positions (1–50)
Clear filtersBuild pipelines that collect, clean and version the evaluation datasets used to measure model quality before release.
Own the indexing and retrieval pipelines that keep the answer engine fresh across the open web.
Build model training, serving and governance features that let enterprises run their own AI workloads on their own data.
Design and run evaluations that measure model capability and safety for enterprise and government customers.
Train, evaluate and release open models, and maintain the libraries the community builds on.
Improve the fraud models scoring billions of payments a year without adding friction for honest buyers.
Build the pipelines behind product analytics for services used by a very large share of the web.
Build semantic and hybrid search features on top of the Elastic Stack for enterprise customers.
Analyse acquisition and activation funnels and turn the results into decisions the team acts on.
Train, evaluate and serve models that power Airbyte's open-source projects at production scale.
Build the ingestion, storage and query layers that every team uses to reason about Airbyte's open-source projects.
Own the lakehouse and streaming infrastructure underneath Amplitude's analytics platform, with an eye on cost and freshness.
Design red-teaming and evaluation methodology for the models behind Anthropic's AI products, and turn results into mitigations.
Design and run large-scale training experiments, then turn the findings into production model releases.
Own the transformation layer and semantic definitions behind reporting on Cursor (Anysphere)'s developer platform.
Build the ingestion, storage and query layers that every team uses to reason about Cursor (Anysphere)'s developer platform.
Run experiments that push the model capabilities behind Atoms's robotics products, and publish or ship what works.
Own the modelling loop behind Automattic's open-source projects — features, training pipelines, offline evaluation and online experiments.
Own the modelling loop behind Buffer's product suite — features, training pipelines, offline evaluation and online experiments.
Train, evaluate and serve models that power Cal.com's open-source projects at production scale.
Optimise inference throughput and cost for the models running behind Chime's financial platform.
Investigate training, alignment and evaluation methods that make the models behind CircleCI's developer platform more useful.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Clay's product suite.
Build the ingestion, storage and query layers that every team uses to reason about Clay's product suite.
Design and run large-scale training experiments, then turn the findings into production model releases.
Optimise inference throughput and cost for the models running behind ClickHouse's data platform.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Close's product suite.
Own the transformation layer and semantic definitions behind reporting on Cloudflare's developer platform.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Coinbase's financial platform.
Own the lakehouse and streaming infrastructure underneath Coinbase's financial platform, with an eye on cost and freshness.
Design and run large-scale training experiments, then turn the findings into production model releases.
Own the lakehouse and streaming infrastructure underneath CrowdStrike's security platform, with an eye on cost and freshness.
Run experiments that push the model capabilities behind Databricks's data platform, and publish or ship what works.
Probe model behaviour for failure modes and build the evaluations that gate releases across Datadog's developer platform.
Train, evaluate and serve models that power Datadog's developer platform at production scale.
Optimise inference throughput and cost for the models running behind dbt Labs's data platform.
Build the ingestion, storage and query layers that every team uses to reason about DigitalOcean's product suite.
Own the lakehouse and streaming infrastructure underneath DoorDash's delivery network, with an eye on cost and freshness.
Own the lakehouse and streaming infrastructure underneath Drata's security platform, with an eye on cost and freshness.
Run experiments that push the model capabilities behind Elastic's developer platform, and publish or ship what works.
Train, evaluate and serve models that power Elastic's developer platform at production scale.
Own the lakehouse and streaming infrastructure underneath Figma's design platform, with an eye on cost and freshness.
Run experiments that push the model capabilities behind Fivetran's data platform, and publish or ship what works.
Probe model behaviour for failure modes and build the evaluations that gate releases across Fivetran's data platform.
Train, evaluate and serve models that power Fly.io's developer platform at production scale.
Own the lakehouse and streaming infrastructure underneath Fly.io's developer platform, with an eye on cost and freshness.
Own the transformation layer and semantic definitions behind reporting on FusionAuth's security platform.
Own the modelling loop behind GitHub's developer platform — features, training pipelines, offline evaluation and online experiments.
Design red-teaming and evaluation methodology for the models behind GitHub's developer platform, and turn results into mitigations.
Probe model behaviour for failure modes and build the evaluations that gate releases across GitLab's developer platform.
No jobs match your search and filters. Try a different keyword or clear the filters.