Senior AI Research Engineer
Emergent
Emergent
Join Emergent, a leader in autonomous coding agents that are revolutionizing software development. We're building systems that generate, test, and deploy production applications directly from plain-language intent, running at a global scale. Our platform has achieved significant milestones, including over $130M in Annualised Revenue and 10M+ users across 190+ countries who have created over 12M applications.
We are solving complex challenges in AI-driven software creation, focusing on correctness, reliability, security, and scalability in production environments. Our team comprises seasoned entrepreneurs, academic achievers, and experts from leading tech companies like Google, Amazon, and Dropbox. We seek ambitious individuals eager for ownership, speed, and making a global impact.
As a Senior AI Research Engineer, you will be instrumental in defining, measuring, and enhancing the capabilities of our autonomous coding agents. Your role will involve translating abstract concepts of agent quality into concrete, verifiable metrics to guide team decisions and advance the field.
You will drive both incremental improvements and significant breakthroughs in agent performance. This position offers a deep dive into agent behavior, evaluation research, and applied training. Key responsibilities include establishing benchmarks for long-horizon coding agents, developing robust evaluation methodologies and datasets, analyzing production data to identify failure modes, and conducting targeted experiments in training, fine-tuning, RL, memory, and prompt optimization. You will operate with substantial autonomy, make critical decisions in complex probabilistic systems, and own outcomes from inception to completion. If you are passionate about rigorous analysis, advancing AI research benchmarks, and applying cutting-edge agent research to real-world applications, this role is for you.
We are seeking candidates with 5 to 8 years of experience in AI, with a strong focus on either training and fine-tuning models or designing rigorous evaluation and measurement systems. Proficiency in Python for research workflows, including training pipelines, evaluation harnesses, data processing, and statistical analysis, is essential. Familiarity with modern AI stacks such as transformers, RLHF/DPO/RL for agents, evaluation frameworks (e.g., Inspect, lm-eval-harness), prompt optimization, judge models, and agent frameworks is required. Go is considered a plus.
This role requires a deep appreciation for data-driven decision-making and the ability to rigorously defend evaluation metrics and results. You should be comfortable navigating subjective and probabilistic systems, understanding concepts like noise floors, confounds, distribution shift, and judge bias. The ideal candidate enjoys diving deep into complex data to uncover subtle failure modes and possesses an intuitive understanding of model behavior, updating their insights based on empirical evidence. We value independent operators with leadership qualities who can clearly articulate their reasoning and drive progress through conviction.
Emergent
IT Consulting