Data Engineer, Evaluation
Build pipelines that collect, clean and version the evaluation datasets used to measure model quality before release.
538 roles across 160 companies — filter by region or category.
Showing 25 of 25 matching positions (1–25)
Clear filtersBuild pipelines that collect, clean and version the evaluation datasets used to measure model quality before release.
Design and run evaluations that measure model capability and safety for enterprise and government customers.
Design red-teaming and evaluation methodology for the models behind Anthropic's AI products, and turn results into mitigations.
Design and run large-scale training experiments, then turn the findings into production model releases.
Optimise inference throughput and cost for the models running behind Chime's financial platform.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Clay's product suite.
Build the ingestion, storage and query layers that every team uses to reason about Clay's product suite.
Own the lakehouse and streaming infrastructure underneath DoorDash's delivery network, with an eye on cost and freshness.
Own the lakehouse and streaming infrastructure underneath Figma's design platform, with an eye on cost and freshness.
Own the transformation layer and semantic definitions behind reporting on FusionAuth's security platform.
Investigate training, alignment and evaluation methods that make the models behind Glean's AI products more useful.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Groq's AI products.
Own the transformation layer and semantic definitions behind reporting on Gusto's employment platform.
Own the transformation layer and semantic definitions behind reporting on Retool's developer platform.
Own the transformation layer and semantic definitions behind reporting on Runway's AI products.
Own the lakehouse and streaming infrastructure underneath Runway's AI products, with an eye on cost and freshness.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Scale AI's AI products.
Investigate training, alignment and evaluation methods that make the models behind Sierra's AI products more useful.
Model user and system behaviour on Runway's AI products and turn the findings into decisions the team acts on.
Build and operate the batch and streaming pipelines that move data across Sierra's AI products.
Answer the recurring commercial and operational questions about Airtable's product suite with clear, reproducible analysis.
Build dashboards and deep dives that show how Anthropic's AI products is actually performing.
Build dashboards and deep dives that show how OpenAI's AI products is actually performing.
Build dashboards and deep dives that show how Runway's AI products is actually performing.
Build models that detect anomalous spending, forecast cash flow and surface savings opportunities across customer accounts.
No jobs match your search and filters. Try a different keyword or clear the filters.