Research Engineer, Alignment
Run experiments that make large language models more honest, harmless and steerable, turning research findings into production safeguards.
538 roles across 160 companies — filter by region or category.
Showing 50 of 177 matching positions (1–50)
Clear filtersRun experiments that make large language models more honest, harmless and steerable, turning research findings into production safeguards.
Build the agentic coding tool developers use in the terminal, the IDE and CI, covering tool use, permissions and long-running sessions.
Design the Claude apps end to end, from conversation surfaces to file handling, for both consumer and enterprise users.
Take custom inference silicon from specification through tapeout and production ramp alongside the model and infrastructure teams.
Coordinate schedules, vendors and bring-up milestones across the custom silicon programme from design freeze to volume production.
Own segmentation, quotas and territory planning for a fast-growing enterprise AI sales organisation.
Grow the cloud marketplace partnerships that put Claude in front of enterprise buyers on AWS and Google Cloud.
Protect model weights and training infrastructure with threat modelling, hardening and detection engineering.
Optimise serving throughput and latency for frontier models across large GPU fleets.
Build the developer-facing API surface — auth, rate limiting, billing and SDKs — used by millions of applications.
Build pipelines that collect, clean and version the evaluation datasets used to measure model quality before release.
Run full enterprise sales cycles for ChatGPT Enterprise and the API platform with large regulated accounts.
Write the API reference, cookbooks and migration guides that developers rely on when shipping with the platform.
Build the labelling and quality tooling that turns raw data into training sets for frontier model developers.
Design and run evaluations that measure model capability and safety for enterprise and government customers.
Run the human review operation behind large annotation programmes, balancing throughput against quality targets.
Sell data and evaluation programmes to enterprises building their own AI capability.
Build the enrichment engine that fans a single record out across dozens of data providers and reconciles the results.
Build proof-of-concept go-to-market workflows for prospects during technical evaluations.
Onboard revenue teams and turn early workflows into durable, expanding accounts.
Build card issuing, authorisation and controls for corporate spend at high transaction volume.
Design the handoff surfaces where engineers translate design files into shipped interfaces.
Build AI writing, search and agent features inside a workspace people use all day.
Investigate intrusions and turn each one into durable detection coverage for Airtable's product suite.
Drive cross-team programmes across Airtable's product suite, keeping dependencies, risk and timelines visible.
Make the build, test and deploy loop fast for everyone working on Airtable's product suite.
Partner with engineering teams on threat models and secure defaults throughout Anthropic's AI products.
Design red-teaming and evaluation methodology for the models behind Anthropic's AI products, and turn results into mitigations.
Design and run the multi-region cloud footprint behind Anthropic's AI products, with cost and resilience both in view.
Design and run large-scale training experiments, then turn the findings into production model releases.
Build detections and response automation covering the infrastructure behind Chime's financial platform.
Run the training and serving infrastructure behind Chime's financial platform, including GPU scheduling, artefacts and rollouts.
Optimise inference throughput and cost for the models running behind Chime's financial platform.
Model warehouse data into tested, documented datasets the whole company trusts when analysing Clay's product suite.
Build detections and response automation covering the infrastructure behind Clay's product suite.
Build the ingestion, storage and query layers that every team uses to reason about Clay's product suite.
Partner with engineering teams on threat models and secure defaults throughout DoorDash's delivery network.
Turn model capabilities into dependable product surfaces across DoorDash's delivery network, balancing latency, cost and quality.
Own infrastructure-as-code, networking and capacity planning for DoorDash's delivery network.
Own the lakehouse and streaming infrastructure underneath DoorDash's delivery network, with an eye on cost and freshness.
Investigate intrusions and turn each one into durable detection coverage for Figma's design platform.
Own the lakehouse and streaming infrastructure underneath Figma's design platform, with an eye on cost and freshness.
Run the large multi-team efforts behind Figma's design platform from kickoff through launch.
Prototype and ship interface work for Figma's design platform, living between design and front-end code.
Own the transformation layer and semantic definitions behind reporting on FusionAuth's security platform.
Build detections and response automation covering the infrastructure behind FusionAuth's security platform.
Investigate intrusions and turn each one into durable detection coverage for Glean's AI products.
Investigate training, alignment and evaluation methods that make the models behind Glean's AI products more useful.
Sit with customers, build the integrations they need on Glean's AI products, and feed the rough edges back to product.
Make the build, test and deploy loop fast for everyone working on Groq's AI products.
No jobs match your search and filters. Try a different keyword or clear the filters.