Software Engineer - Infrastructure
Emergent
Emergent
Join Emergent, a pioneering company revolutionizing software development with autonomous coding agents. Our AI-driven platform generates, tests, and deploys production applications, operating at a global scale. We've achieved significant milestones, including over $100M in ARR and a user base exceeding 10 million worldwide.
We are at the forefront of solving complex AI software creation challenges related to correctness, reliability, security, and scalability. Our team comprises accomplished individuals, including repeat founders, Olympiad medalists, and alumni from top institutions, with extensive experience from leading tech companies. We are seeking passionate builders who value ownership, speed, and global impact.
This role offers an exciting opportunity to work on cutting-edge infrastructure and contribute to the future of software development.
Maintain the stability of our distributed microservices platform, integrated with Kubernetes and cloud providers (GCP, AWS). Manage Kubernetes workloads using ArgoCD for GitOps, handling deployments, monitoring, and troubleshooting. Debug and resolve intricate Kubernetes issues across various clusters. Oversee CDN and edge infrastructure (Cloudflare) for optimal performance, caching, and traffic management. Automate infrastructure lifecycle operations and associated workflows. Own and enhance the observability stack, including Grafana, Loki, Prometheus, and New Relic, for monitoring, alerting, and distributed tracing. Participate in on-call rotations, manage incident response, and conduct root cause analysis. Proactively identify and mitigate potential reliability risks. Support the platform powering AI agent workloads, encompassing job scheduling, trajectory tracking, environment provisioning, deployments, and cost attribution. Develop Kubernetes controllers and operators to expand platform capabilities for agent orchestration. Collaborate closely with product and backend teams to ensure platform scalability and reliability. Build internal tools, automate workflows, and integrate systems to boost team productivity. Stay abreast of the latest Kubernetes releases, CNCF ecosystem advancements, and cloud-native best practices.
We are looking for candidates with at least 3 years of experience in software or platform engineering, specifically with production systems. Proficiency in Go or Python is essential, with the ability to write production-ready code in at least one daily. Demonstrated hands-on experience building and deploying services on Kubernetes, beyond just managing YAML configurations. Familiarity with GitOps tooling such as ArgoCD, Flux, or similar solutions. Strong understanding of systems fundamentals, including networking (TCP/IP, HTTP, load balancing, DNS, TLS) and Linux/OS concepts (process management, filesystem, memory, systemd). Experience with relational databases (PostgreSQL, MySQL) covering indexing, query optimization, replication, and backup procedures. Familiarity with NoSQL databases (MongoDB, DynamoDB, Redis) and caching solutions (Redis, Memcached). Hands-on experience with message queues and streaming platforms like Kafka, SQS, or RabbitMQ. Comfort with the CNCF ecosystem, including tools like Helm and Kustomize. Proficiency with at least one observability stack (Grafana/Prometheus/Loki, New Relic, Datadog). Experience with cloud platforms such as GCP and/or AWS, including managed Kubernetes services.
Familiarity with CDN/edge platforms (Cloudflare, Cloudfront) is beneficial. Experience building Kubernetes Operators and tuning Kubernetes core components is a plus. Knowledge of AI/LLM infrastructure and CI/CD pipelines is advantageous. Infrastructure as Code experience (Terraform, Pulumi) and previous work on large-scale distributed systems are highly desirable. We value individuals who can context-switch effectively, enjoy solving complex infrastructure challenges, possess strong debugging skills, and communicate clearly within a fast-moving team.
Emergent
Technology