Data Scientist
Emergent
Emergent
Emergent is revolutionizing software development with autonomous AI coding agents that generate, test, and deploy production applications from simple language prompts. Our scalable systems are already used to build millions of real-world applications. We've achieved significant growth, crossing $130M in Annualised Revenue and attracting over 10M users globally, who have built more than 12M applications on our platform. Our success is supported by leading investors like Creaegis, Khosla Ventures, and SoftBank.
We tackle the most challenging aspects of AI-driven software creation: ensuring correctness, reliability, security, and scalability in production environments. Our team comprises seasoned entrepreneurs, top academic achievers, and experts from tech giants like Google and Amazon. We are seeking passionate builders who thrive on ownership, speed, and making a global impact.
This role is crucial for leveraging data to drive every aspect of our growth, from user activation and conversion to understanding agent-built application costs and identifying misuse of our free tier. You will work with a unique dataset that includes vast amounts of unstructured signal, such as agent trajectories, support tickets, user feedback, and natural-language prompts, alongside standard event and revenue data. Your primary responsibility will be to extract rich insights from this text data.
You will own the entire analytics loop: defining what we measure, developing methods to prove insights, and interpreting the implications of the data. This involves transforming agent trajectories, support tickets, and user prompts into structured, queryable signals through advanced pipelines. You will identify early indicators of product issues, churn risk, or fraud, and ensure these are communicated to the relevant teams. Additionally, you will build predictive models for conversion, retention, and churn, integrating these signals into product and growth workflows. A key part of this role includes owning marketing attribution and MMM, developing models to understand what drives user acquisition and paid conversions, especially in environments with limited per-user attribution.
You will also lead product and growth analytics across our self-serve funnel, on both web and mobile platforms, covering activation, engagement, and conversion. This includes designing and rigorously analyzing A/B tests and growth experiments, considering factors like power, novelty effects, interference, and causal inference. You will conduct clustering analysis on agent trajectories to identify common user building patterns and failure modes, providing product teams with a non-labeled taxonomy that predicts churn. Furthermore, you will model the user 'aha moment,' using text-derived features from initial interactions to optimize onboarding for long-term retention. Building a gross-margin model to attribute LLM and compute costs to individual apps and cohorts will also be a core duty, enabling product teams to understand net-positive segments. Finally, you will investigate and resolve complex issues, such as untangling fraud rings and determining whether they represent an enforcement or pricing problem, backing your recommendations with robust data.
We are seeking candidates with 2 to 5 years of experience in data science or applied machine learning, with a strong emphasis on product analytics, growth, or user behavior. A deep proficiency in SQL and extensive experience working with large-scale, event-level behavioral data are essential. Candidates should possess solid classical machine learning foundations, including expertise in clustering (k-means, HDBSCAN, hierarchical), embeddings, vector similarity, dimensionality reduction (UMAP/PCA), and classification, with a clear understanding of their appropriate applications.
Genuine skill in deriving insights from unstructured natural language data, such as LLM traces, logs, tickets, and free text, and translating it into predictive, queryable signals is highly valued. Familiarity with topic modeling and trace-clustering approaches like Clio-style summarize-then-embed pipelines, BERTopic, or c-TF-IDF is a significant advantage. Proficiency in Python and the standard data science stack (pandas, scikit-learn, statsmodels, NumPy) is required. Competence in data engineering, including designing and shipping ETL processes and data models (using dbt or equivalent), is also expected.
Candidates must be experienced in designing and analyzing experiments, with a strong grasp of sample sizing, statistical power, significance, novelty effects, test interference, and causal methods. The ability to move quickly and deeply, delivering analysis efficiently through developed intuition and rigorous self-testing, is critical. You should be adept at investigating top-line metrics to uncover underlying confounders, questioning the assumptions behind metrics, and always articulating the support, limitations, and next steps for any data presented. While leveraging AI agents to enhance productivity is encouraged, all AI-assisted results must be treated as drafts requiring validation. A profound curiosity for user behavior and an instinct for drivers of growth, retention, and abuse are paramount. The role demands fluidity across exploratory analysis, ML modeling, hypothesis testing, and translating findings into strategic recommendations that drive decision-making.
Emergent
AI / Machine Learning