Software Solutions Engineer
NVIDIA
NVIDIA
Join NVIDIA's AI Enterprise team as a Software Solutions Engineer, supporting customer deployments across cloud and datacenter environments. This dynamic role involves end-to-end customer issue resolution and the development of new software features, automation, and diagnostics to enhance product readiness and support scalability.
You'll engage with cutting-edge compute and cloud-native technologies, including container platforms, enterprise system software, and GPU-accelerated AI frameworks. This customer-facing position requires close collaboration with clients and internal engineering teams to understand, diagnose, and resolve complex issues, driving improvements and fixes.
Success hinges on strong debugging capabilities, clear communication, and a commitment to owning technical escalations from start to finish.
Develop and maintain product features and deployment assets for AI Enterprise, including scripts, configuration guides, Kubernetes manifests, and reproducible test cases.
Create and manage Python-based automation and tooling to boost NVIDIA AI Enterprise deployment reliability on NGC and container orchestrators like Kubernetes.
Contribute code-level fixes and patches in partnership with engineering to resolve customer-impacting issues and enhance product quality.
Provide support for enterprise customers deploying NVIDIA AI Enterprise in datacenter and cloud service provider environments, focusing on Kubernetes-based and containerized AI platforms.
Take full ownership of customer issues, from reproduction in lab/cloud environments to collecting diagnostics, providing workarounds, and collaborating with engineering on solutions.
Produce detailed bug reports and feature requests with clear reproduction steps, environment specifics, impact analysis, and supporting evidence.
Develop both customer-facing and internal documentation, such as knowledge base articles and runbooks, to accelerate time-to-value and minimize recurring problems.
Participate in an on-call rotation, providing engineering assistance for critical customer outages one weekend per month.
A Bachelor of Science in Computer Science, Electrical Engineering, Computer Engineering, or a related field, or equivalent practical experience.
Minimum of 5 years in system software development and troubleshooting, with a preference for customer-facing experience.
Proficiency in computer science fundamentals and programming/scripting, particularly Python, with Bash knowledge and Go/C++ being beneficial for automation and diagnostic tool development.
Strong troubleshooting acumen in networking, concurrency, and OS concepts, coupled with a methodical approach to isolating issues across application, platform, and infrastructure layers.
In-depth understanding of at least two core areas: data centers/servers, distributed systems, virtualization, deep learning frameworks, containers (Docker/Kubernetes), hybrid cloud (AWS/Azure/GCP), or CI/CD practices for robust deployments.
Familiarity with GPU-accelerated AI/ML stacks and production model deployment/serving, including NGC containers, CUDA concepts, and inference servers like Triton.
Extensive knowledge of Linux and comfort troubleshooting in production Linux environments; Windows proficiency is a plus.
Professional-level communication and interpersonal skills, driven by a passion for problem-solving.
Nvidia
IT Consulting