Experience: 5–7+ Years
Domain: Azure Cloud Infrastructure | Software Engineering | Diagnostics | Quality Engineering | Live-Site Operations

Ideal Candidate Profile:
5–7+ years | C#/Python/PowerShell | Azure/Cloud Infrastructure | Distributed Systems | Live-Site & RCA | Hardware Diagnostics | Telemetry & Observability | Git/CI/CD | AI/HPC preferred

Pan India – remote role
Experience- 4+ works
Skills – C# .net, Azure Infrastructure


Role Overview
We are looking for an experienced Azure Software Design Engineer with 5–7 years of relevant experience in software engineering, cloud infrastructure, distributed systems, or production engineering.
The role focuses on developing and supporting Azure infrastructure software, hardware diagnostics, SKU qualification, telemetry/observability, and live-site operations. The ideal candidate will have strong debugging and root-cause analysis skills and be comfortable troubleshooting complex software, hardware, networking, and infrastructure issues in large-scale production environments.
Key Responsibilities

  • Develop and maintain backend services, diagnostics, and infrastructure tools using C#, .NET, Python, PowerShell, C++, or Rust.
  • Execute and monitor hardware SKU qualification and host networking pipelines, investigate failures, and rerun failed scenarios.
  • Perform live-site incident triage, debugging, fault isolation, RCA, mitigation, and recovery.
  • Analyze flaky tests and recurring failures and drive improvements in test coverage, reliability, scale, and execution efficiency.
  • Monitor and troubleshoot HDTest and hardware diagnostic setup/execution failures.
  • Develop and support hardware health monitoring, fault attribution, diagnostics, certification, remediation, and recovery workflows.
  • Enhance telemetry, metrics, observability, monitoring, dashboards, and operational documentation.
  • Support hardware-adjacent systems involving BMC, firmware, server management, provisioning, inventory, deployment, and device operations.
  • Develop and maintain CI/CD pipelines using Azure DevOps, GitHub Actions, or equivalent technologies.
  • Manage source code through Git, branching, pull requests, code reviews, and release workflows.
  • Develop automation using C#, Python, and PowerShell to improve operational efficiency.
  • Collaborate with software, hardware, networking, quality, PM, and datacenter operations teams.

Required Qualifications

  • 5–7 years of relevant professional experience in software engineering, cloud infrastructure, systems engineering, SRE, production engineering, or infrastructure diagnostics.
  • Strong software engineering fundamentals in one or more of C#, C++, Rust, Python, or PowerShell.
  • Proven experience in debugging, troubleshooting, fault isolation, and root-cause analysis of distributed production systems.
  • Experience supporting large-scale cloud infrastructure, datacenter platforms, hardware management, or distributed systems.
  • Hands-on experience with live-site operations, incident management, triage, mitigation, and recovery workflows.
  • Experience with telemetry, logging, diagnostics, monitoring, observability, and fault-attribution systems.
  • Experience with hardware-adjacent systems such as BMC, firmware, server diagnostics, inventory, provisioning, or deployment.
  • Strong knowledge of Git and CI/CD workflows.
  • Ability to independently investigate complex problems in ambiguous environments and drive them through resolution.
  • Strong communication and cross-functional collaboration skills.

Preferred Qualifications

  • Experience with distributed systems, orchestration engines, stateful workflows, and scalable control planes.
  • Experience with Linux-based development and diagnostics.
  • Experience with hardware diagnostics, health monitoring, certification, remediation, or recovery solutions.
  • Knowledge of servers, networking, storage, rack-level architecture, and accelerator infrastructure.
  • Experience with fleet management, firmware deployment, provisioning, inventory, or device operations.
  • Experience with Infrastructure as Code using Terraform, ARM templates, or Bicep.
  • Experience with AI/HPC infrastructure, GPU clusters, Bare Metal Instances, or hyperscale compute environments.
  • Knowledge of RDMA, InfiniBand, and accelerator-focused infrastructure is an advantage.
  • Understanding of reliability engineering, MTTR optimization, resiliency, and operational excellence.

Azure Infrastructure Focus
Experience in the following areas is highly desirable:

  • Hardware health monitoring and fault attribution
  • AI/HPC SKU onboarding and qualification
  • GB300, VR200, MI455x, Maia300, or similar accelerator platforms
  • Bare Metal Instance diagnostics and validation
  • Device recovery and automated remediation
  • Inventory, provisioning, firmware deployment, and certification
  • Integration with Titan, SCHIE, OneMOS, APlat, and Azure infrastructure services

Preferred Certifications

  • Microsoft Azure Developer Associate
  • Microsoft DevOps Engineer Expert
  • Microsoft Azure Solutions Architect Expert
  • Relevant cloud, Linux, networking, or infrastructure certifications

Apply for this position

Allowed Type(s): .pdf, .doc, .docx