Responsibilities: 

  • Infrastructure Deployment & Support : Deploy server and network infrastructure equipment at cloud scale in collaboration with data center , software teams and platform engineering teams 
  • Hardware Operations & Server Lifecycle Management : Engineering guidance for Hardware Troubleshooting and Lifecycle support for Dell PowerEdge and co-ordinate with OEM Vendors 
  • Firmware, BIOS, Driver and Patch Management : Hypervisor, Firmware and Node upgrades as managed services to support Infrastructure Patching 
  • Day2 Operations & Live-site support : Triage hardware, compute, capacity, deployment, and platform health issues. 
  • Network and Rack-Level Coordination : Coordinate compute deployment dependencies with network teams during rack turn-up and infrastructure validation 
  • Incident Management and Operational Governance : Use incident management tools such as IcM or equivalent systems to track ownership, status, mitigation, and closure 
  • Documentation and Knowledge Management : Create and maintain technical documentation, SOPs, TSGs, runbooks, checklists, and operational handoff documents. 
  • Automation and Continuous Improvement : Create scripts for validation, patching checks, health checks, inventory collection, and operational reporting.  

Qualifications: 

Education: 

  • Bachelor’s or Master’s Degree in Computer Science, Information Technology, or a related field. 

Technical Experience: 

  • Around 5 years of experience in compute infrastructure, data center operations, systems administration, cloud operations, or infrastructure buildout. 
  • Assist in server rack bring-up, hardware validation, host discovery, power-on validation, and compute node readiness checks 
  • Experience using Azure DevOps, pipelines, work item tracking, release planning, or infrastructure rollout processes 
  • Hands-on experience troubleshooting real-world compute hardware issues across Dell PowerEdge platform in large-scale cloud or data center environments. 
  • Execute firmware, BIOS, driver, RAID, and BMC update activities across compute nodes 
  • Support LSI and CRI handling by gathering logs, validating impact, coordinating mitigation, and documenting recovery steps 
  • Validate server-side NIC status, cabling, link state, SFP compatibility, and port connectivity 

Specialized Skills: 

  • Experience with : Azure infrastructure fundamentals, VM lifecycle management, Linux server administration, Server hardware health checks, 
  • Certifications  : Dell administration certificate, Microsoft Azure fundamentals, ITIL Foundation, Linux Foundation or Red Hat Certification  

Soft Skills: 

  • Strong analytical and troubleshooting ability with organizational skills to focus on meeting deadlines and achieving project goals 
  • Ability to collaborate effectively with cross-functional teams 

Apply for this position

Allowed Type(s): .pdf, .doc, .docx