ITS Systems Engineer, Corporate Infrastructure Services, IT

Amazon

4+ yrs Bengaluru Full Time Hybrid (office + remote)
Amazon logo
Posted : yesterday
Actively hiring

Job description

Join our Corporate Infrastructure Services team, responsible for building and operating the global systems that power Amazon's corporate offices. This includes networks, compute platforms, AV systems, and operational tooling essential for every Amazonian.

Our Systems Engineering function drives operational standards, fleet governance, observability, and engineering practices to ensure reliable, secure, and cost-effective worldwide infrastructure. We are seeking a Systems Engineer focused on fleet management and operational excellence to maintain infrastructure health, compliance, and lifecycle management.

This role blends hands-on systems engineering with a strategic, fleet-level operational perspective. You will tackle complex systems challenges across hardware, software, networking, and cloud environments. The position involves developing automation to scale operations, replacing manual tasks with robust, repeatable solutions. You will also create and maintain SOPs, runbooks, and documentation for consistent global operations and identify patterns impacting fleet health, driving improvements for measurable results.

Embrace autonomy, taking ownership of issues that span multiple domains. Proficiency in operating systems, networking, compute hardware, cloud platforms, and monitoring tools is key. You will leverage AWS services (EC2, Lambda, Systems Manager, CloudWatch, DynamoDB, S3, Fleet Manager) and scripting languages (Python, PowerShell, Bash) to build efficient and scalable fleet operations tooling.

Responsibilities

Take ownership of the operational health and compliance of infrastructure fleets, ensuring accurate inventory, tracking lifecycle status, and adhering to fleet governance standards.

Diagnose and resolve complex systems issues across hardware, software, networking, and operating environments, driving to root cause identification.

Develop automation solutions for scaling fleet operations, including device discovery, compliance scanning, firmware tracking, health monitoring, and lifecycle reporting.

Create, review, and enhance SOPs, runbooks, and documentation to guarantee consistent and repeatable operations across all sites and regions.

Identify trends impacting fleet performance, reliability, availability, or compliance, and implement scalable automation to address them.

Champion operational excellence initiatives to achieve measurable improvements in fleet health, compliance posture, and operational efficiency.

Provide cross-domain insights to software, hardware, networking, and security engineers regarding component interactions within the system.

Participate in on-call rotations, diagnosing and resolving operational issues across the entire infrastructure estate.

Mentor junior engineers, guiding them in systems architecture, operational best practices, and fleet management disciplines.

Contribute valuable operational and fleet perspectives to team design, scoping, and prioritization discussions.

Qualifications

Possess at least 4 years of experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration.

Have 5+ years of hands-on experience with Linux operating systems.

Demonstrate 5+ years of comprehensive systems engineering experience.

Hold a Bachelor's degree in Systems Engineering, Computer Science, or a related field, or possess equivalent relevant work experience.

Exhibit experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration.

Demonstrate proficiency in at least one of the following scripting languages: Python, Java, Perl, PHP, Ruby, Bash, Shell, or equivalent.

Familiarity with TCP/IP and networking protocols like HTTP and DNS is advantageous.

Experience in designing and developing scripts for automating operational tasks, including reviewing scripting changes for maintainability, scalability, and security, is preferred.

Experience working within a 24/7 production environment is a plus.

Knowledge of service-oriented architecture and web services is beneficial.

Essential Skills

LinuxPythonBashSystems EngineeringSite Reliability Engineering (SRE)DevOpsSecurity AdministrationNetwork AdministrationAWSEC2LambdaSystems ManagerCloudWatchDynamoDBS3Fleet Manager

Good to Have

TCP/IPHTTPDNSService-Oriented ArchitectureWeb ServicesJavaPerlPHPRubyShell

Highlights

  • Actively hiring

More Details

RoleITS Systems Engineer, Corporate Infrastructure Services, IT
IndustryInformation Technology & Services
DepartmentEngineering, System Administration, Information Technology
Employment TypeFull Time, Hybrid (office + remote)

About the Company

Amazon logo

Amazon

Information Technology & Services

ITS Systems Engineer, Corporate Infrastructure Services, IT at Amazon | SkillMX | SkillMX