Research Engineer, Alignment
Run experiments that make large language models more honest, harmless and steerable, turning research findings into production safeguards.
538 roles across 160 companies — filter by region or category.
Showing 50 of 318 matching positions (1–50)
Clear filtersRun experiments that make large language models more honest, harmless and steerable, turning research findings into production safeguards.
Build the agentic coding tool developers use in the terminal, the IDE and CI, covering tool use, permissions and long-running sessions.
Design the Claude apps end to end, from conversation surfaces to file handling, for both consumer and enterprise users.
Take custom inference silicon from specification through tapeout and production ramp alongside the model and infrastructure teams.
Coordinate schedules, vendors and bring-up milestones across the custom silicon programme from design freeze to volume production.
Own segmentation, quotas and territory planning for a fast-growing enterprise AI sales organisation.
Grow the cloud marketplace partnerships that put Claude in front of enterprise buyers on AWS and Google Cloud.
Protect model weights and training infrastructure with threat modelling, hardening and detection engineering.
Optimise serving throughput and latency for frontier models across large GPU fleets.
Build the developer-facing API surface — auth, rate limiting, billing and SDKs — used by millions of applications.
Build pipelines that collect, clean and version the evaluation datasets used to measure model quality before release.
Run full enterprise sales cycles for ChatGPT Enterprise and the API platform with large regulated accounts.
Write the API reference, cookbooks and migration guides that developers rely on when shipping with the platform.
Squeeze latency and cost out of the serving stack behind a consumer answer engine handling very high query volume.
Build multi-step agents that browse, reason over sources and complete tasks on behalf of users.
Own the indexing and retrieval pipelines that keep the answer engine fresh across the open web.
Ship experiments across onboarding, sharing and subscription flows to move activation and retention.
Sell Perplexity Enterprise to mid-market teams, running discovery through to close.
Own paid, lifecycle and partnership channels driving qualified pipeline for the enterprise product.
Build the query and storage layer of the lakehouse platform, working across Spark, Delta Lake and cloud object storage.
Join a platform team building distributed data systems, with structured mentoring through your first year.
Build model training, serving and governance features that let enterprises run their own AI workloads on their own data.
Design lakehouse architectures with enterprise customers and prove them out in technical evaluations.
Own governance, lineage and access control across data and AI assets for large regulated customers.
Document pipelines, SQL warehousing and workflow orchestration for practitioners across cloud providers.
Build the labelling and quality tooling that turns raw data into training sets for frontier model developers.
Design and run evaluations that measure model capability and safety for enterprise and government customers.
Run the human review operation behind large annotation programmes, balancing throughput against quality targets.
Sell data and evaluation programmes to enterprises building their own AI capability.
Build the AI agent that scaffolds, edits and deploys full applications from a natural-language brief.
Run the container and sandbox fleet that gives every user an instant, isolated development environment.
Design an IDE that stays approachable for first-time programmers without frustrating professionals.
Build the enrichment engine that fans a single record out across dozens of data providers and reconciles the results.
Build proof-of-concept go-to-market workflows for prospects during technical evaluations.
Onboard revenue teams and turn early workflows into durable, expanding accounts.
Build the access, gateway and browser isolation controls enterprises use to retire their VPNs.
Build ingestion and query systems handling trillions of log events with predictable cost.
Build card issuing, authorisation and controls for corporate spend at high transaction volume.
Design the handoff surfaces where engineers translate design files into shipped interfaces.
Build AI writing, search and agent features inside a workspace people use all day.
Investigate intrusions and turn each one into durable detection coverage for Airtable's product suite.
Drive cross-team programmes across Airtable's product suite, keeping dependencies, risk and timelines visible.
Make the build, test and deploy loop fast for everyone working on Airtable's product suite.
Partner with engineering teams on threat models and secure defaults throughout Anthropic's AI products.
Design red-teaming and evaluation methodology for the models behind Anthropic's AI products, and turn results into mitigations.
Design and run the multi-region cloud footprint behind Anthropic's AI products, with cost and resilience both in view.
Design and run large-scale training experiments, then turn the findings into production model releases.
Build detections and response automation covering the infrastructure behind Chime's financial platform.
Run the training and serving infrastructure behind Chime's financial platform, including GPU scheduling, artefacts and rollouts.
Optimise inference throughput and cost for the models running behind Chime's financial platform.
No jobs match your search and filters. Try a different keyword or clear the filters.