Senior Staff Engineer - Site Reliability Engineering

Altimetrik

6–10 yrs Chennai Full Time Hybrid (office + remote)
Altimetrik logo
Posted : 1 week ago
Actively hiring

Job description

We are seeking a Senior Staff Engineer specializing in Site Reliability Engineering with a strong background in the automotive domain. The ideal candidate will bring 6-10 years of advanced experience in core SRE principles and practices. This role emphasizes building and maintaining highly scalable, reliable, and performant systems.

This position requires deep expertise in cloud platforms like AWS and GCP, leveraging their capabilities for enhanced system resilience. A solid understanding of developing microservices and APIs using Spring Boot and Java is essential. Proficiency in Python and JavaScript for scalable application development and data visualization is also key.

Further versatility is gained through familiarity with languages such as C, C++, and Ruby. The role demands strong data management skills with both SQL and NoSQL databases. Hands-on experience with containerization using Docker and orchestration with Terraform is expected. Familiarity with monitoring tools like Splunk, Nagios, and Prometheus, along with incident management via PagerDuty, is crucial for maintaining system health and availability. A proactive approach to problem-solving, with a focus on automation and continuous improvement, is highly valued.

Responsibilities

Architect, implement, and maintain robust SRE solutions on cloud platforms (AWS, GCP). Develop and deploy microservices and APIs utilizing Spring Boot and Java. Build scalable applications and data-driven dashboards using Python and JavaScript. Manage and optimize data storage and retrieval with SQL and NoSQL databases. Implement and manage containerized applications using Docker and orchestration tools like Terraform. Establish and refine monitoring systems using Splunk, Nagios, and Prometheus to ensure optimal performance and availability. Lead incident response efforts using PagerDuty, focusing on rapid resolution and root cause analysis. Drive automation initiatives to improve operational efficiency and reduce manual toil. Continuously seek opportunities for system improvements and promote best practices in reliability engineering.

Qualifications

Minimum of 6 years of professional experience in Site Reliability Engineering or a related field, with a preference for candidates with 10 years of experience. Advanced proficiency in cloud platforms, specifically AWS and GCP, including cloud architecture and scalability. Expertise in developing microservices and APIs using Spring Boot and Java. Skilled in Python and JavaScript for application development and data presentation. Familiarity with C, C++, and Ruby is advantageous. Strong understanding and practical experience with SQL and NoSQL databases. Hands-on experience with Docker for containerization. Proficiency in using Terraform for infrastructure as code and deployment. Experience with monitoring tools such as Splunk, Nagios, and Prometheus. Demonstrated experience with incident management tools like PagerDuty. Solid grasp of automation principles and practices. Bachelor of Engineering (B.E.) in Computer Science or a related field is required. Master of Technology (M.Tech) in Computer Science/Information Technology is required. AWS Certified Solutions Architect – Professional and Google Professional Cloud Architect certifications are highly preferred.

Essential Skills

AWSGCPSpring BootJavaPythonJavaScriptSQLNoSQLDockerTerraformSplunkNagiosPrometheusPagerDutyAPIMicroservices

Good to Have

CC++RubyDashboardsAutomationReliability Engineering

Highlights

  • Actively hiring

More Details

RoleSenior Staff Engineer - Site Reliability Engineering
DepartmentSoftware Development
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Altimetrik logo

Altimetrik

IT Consulting