Research Engineer, Alignment
Run experiments that make large language models more honest, harmless and steerable, turning research findings into production safeguards.
1638 roles across 160 companies — filter by region or category.
Showing 50 of 309 matching positions (1–50)
Clear filtersRun experiments that make large language models more honest, harmless and steerable, turning research findings into production safeguards.
Build the agentic coding tool developers use in the terminal, the IDE and CI, covering tool use, permissions and long-running sessions.
Design the Claude apps end to end, from conversation surfaces to file handling, for both consumer and enterprise users.
Take custom inference silicon from specification through tapeout and production ramp alongside the model and infrastructure teams.
Coordinate schedules, vendors and bring-up milestones across the custom silicon programme from design freeze to volume production.
Own segmentation, quotas and territory planning for a fast-growing enterprise AI sales organisation.
Grow the cloud marketplace partnerships that put Claude in front of enterprise buyers on AWS and Google Cloud.
Protect model weights and training infrastructure with threat modelling, hardening and detection engineering.
Optimise serving throughput and latency for frontier models across large GPU fleets.
Build the developer-facing API surface — auth, rate limiting, billing and SDKs — used by millions of applications.
Build pipelines that collect, clean and version the evaluation datasets used to measure model quality before release.
Run full enterprise sales cycles for ChatGPT Enterprise and the API platform with large regulated accounts.
Write the API reference, cookbooks and migration guides that developers rely on when shipping with the platform.
Build the AI-native editing experience — completions, multi-file edits and review flows — used by professional engineering teams.
Run the GPU and indexing infrastructure behind low-latency code intelligence at very large repository scale.
Define how background coding agents plan, execute and hand work back to developers.
Own the self-serve funnel from install to paid team plan, instrumenting and testing every step.
Build the labelling and quality tooling that turns raw data into training sets for frontier model developers.
Design and run evaluations that measure model capability and safety for enterprise and government customers.
Run the human review operation behind large annotation programmes, balancing throughput against quality targets.
Sell data and evaluation programmes to enterprises building their own AI capability.
Build the enrichment engine that fans a single record out across dozens of data providers and reconciles the results.
Build proof-of-concept go-to-market workflows for prospects during technical evaluations.
Onboard revenue teams and turn early workflows into durable, expanding accounts.
Build card issuing, authorisation and controls for corporate spend at high transaction volume.
Investigate suspicious account activity and tune the rules behind fraud and AML detection.
Analyse acquisition and activation funnels and turn the results into decisions the team acts on.
Build and maintain the integrations that connect financial apps to thousands of institutions.
Design the handoff surfaces where engineers translate design files into shipped interfaces.
Build AI writing, search and agent features inside a workspace people use all day.
Turn model capabilities into dependable product surfaces across Abridge's healthcare platform, balancing latency, cost and quality.
Design and run the multi-region cloud footprint behind Abridge's healthcare platform, with cost and resilience both in view.
Investigate intrusions and turn each one into durable detection coverage for Abridge's healthcare platform.
Investigate intrusions and turn each one into durable detection coverage for Airtable's product suite.
Drive cross-team programmes across Airtable's product suite, keeping dependencies, risk and timelines visible.
Make the build, test and deploy loop fast for everyone working on Airtable's product suite.
Partner with engineering teams on threat models and secure defaults throughout Anthropic's AI products.
Design red-teaming and evaluation methodology for the models behind Anthropic's AI products, and turn results into mitigations.
Design and run the multi-region cloud footprint behind Anthropic's AI products, with cost and resilience both in view.
Design and run large-scale training experiments, then turn the findings into production model releases.
Build agentic workflows and tool-calling systems on top of Cursor (Anysphere)'s developer platform, and own their accuracy in production.
Own the transformation layer and semantic definitions behind reporting on Cursor (Anysphere)'s developer platform.
Build the ingestion, storage and query layers that every team uses to reason about Cursor (Anysphere)'s developer platform.
Run experiments that push the model capabilities behind Atoms's robotics products, and publish or ship what works.
Own Kubernetes, deployment tooling and paved-road abstractions underneath Atoms's robotics products.
Prototype and ship interface work for Brex's financial platform, living between design and front-end code.
Run the large multi-team efforts behind Brex's financial platform from kickoff through launch.
Deliver working implementations of Brex's financial platform inside customer environments.
Turn model capabilities into dependable product surfaces across Bucket Robotics's robotics products, balancing latency, cost and quality.
Make the build, test and deploy loop fast for everyone working on Bucket Robotics's robotics products.
No jobs match your search and filters. Try a different keyword or clear the filters.