Sandisk

Sandisk

Posted via SmartRecruiters

HPC & Cloud Engineer

Posted Aug 5, 2026

Role at a glance

Salary
Not Disclosed
Location
Bengaluru, KA, India
Work arrangement
Hybrid
Employment
Full-time
Education
Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or related field.

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Role Summary

AI-generated

This role builds and operates large-scale HPC environments across cloud and on-premises platforms for AI/ML and data-intensive workloads. It supports research and engineering teams through infrastructure automation, AI/data pipelines, platform monitoring, and secure, reliable computing services.

What You'll Do

  • Design hybrid-cloud and multi-cloud architectures for HPC workloads.
  • Develop Infrastructure as Code and CI/CD pipelines for infrastructure and platform deployments.
  • Automate cluster provisioning, configuration management, monitoring, and patch management.
  • Design and implement scalable AI/ML data pipelines and support distributed AI training and inference workloads.
  • Implement observability solutions and troubleshoot performance bottlenecks across compute, storage, network, and AI frameworks.
  • Design, deploy, and manage large-scale HPC clusters and administer compute, storage, networking, and GPU resources.

Generated from the employer's posting. Verify important details before applying.

View full posting

Qualifications

DevOps and cloud infrastructure management experience; Linux system administration; HPC clusters and distributed computing; Python, Bash, or Go; Terraform, Ansible, Git, Jenkins/GitHub Actions; Docker, Kubernetes, Singularity/Apptainer; AI/ML frameworks; GPU technologies; AWS, Azure, or GCP architecture and operations; hybrid cloud and HPC workload migration.

Required

  • 5+ years of experience in DevOps and Cloud infrastructure management
  • Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or related field
  • Strong experience with Linux system administration (RHEL, Rocky Linux, Ubuntu)
  • Experience managing HPC clusters and distributed computing environments
  • Proficiency in Python, Bash, or Go
  • Hands-on experience with Terraform, Ansible, Git, Jenkins/GitHub Actions
  • Experience with container technologies: Docker, Kubernetes, Singularity/Apptainer
  • Knowledge of AI/ML frameworks: TensorFlow, PyTorch, Ray, Spark

Original job description

Content provided by the employer

Company Description

Sandisk understands how people and businesses consume data and we relentlessly innovate to deliver solutions that enable today’s needs and tomorrow’s next big ideas. With a rich history of groundbreaking innovations in Flash and advanced memory technologies, our solutions have become the beating heart of the digital world we’re living in and that we have the power to shape.

Sandisk meets people and businesses at the intersection of their aspirations and the moment, enabling them to keep moving and pushing possibility forward. We do this through the balance of our powerhouse manufacturing capabilities and our industry-leading portfolio of products that are recognized globally for innovation, performance and quality.

Sandisk has two facilities recognized by the World Economic Forum as part of the Global Lighthouse Network for advanced 4IR innovations. These facilities were also recognized as Sustainability Lighthouses for breakthroughs in efficient operations. With our global reach, we ensure the global supply chain has access to the Flash memory it needs to keep our world moving forward.

Job Description

Cloud Architecture & Operations

  • Build and operate HPC environments on cloud platforms such as:
    • Amazon Web Services (AWS)
    • Microsoft Azure
    • Google Cloud Platform
  • Design hybrid-cloud and multi-cloud architectures for HPC workloads.

  • Implement cloud-native storage, networking, security, and disaster recovery solutions.

Infrastructure Automation & DevOps

  • Develop Infrastructure as Code (IaC) using:
    • Terraform
    • CloudFormation
    • Ansible

    • Python code

  • Build CI/CD pipelines for infrastructure and platform deployments.
  • Automate cluster provisioning, configuration management, monitoring, and patch management.
  • Develop self-service provisioning frameworks for research and engineering teams.

AI & Data Engineering

  • Design and implement scalable AI/ML data pipelines.
  • Build data ingestion, transformation, and orchestration frameworks.
  • Support distributed AI training and inference workloads.
  • Optimize GPU utilization for deep learning applications.
  • Collaborate with Data Scientists and ML Engineers to deploy production AI solutions.

Platform Monitoring & Reliability

  • Implement observability solutions using: Prometheus, Grafana, ELK Stack, OpenTelemetry
  • Monitor system performance, capacity planning, and SLA compliance.
  • Troubleshoot performance bottlenecks across compute, storage, network, and AI frameworks.

HPC Infrastructure Engineering

  • Design, deploy, and manage large-scale HPC clusters across on-premises and cloud environments.
  • Administer compute, storage, networking, and GPU resources for AI/ML and data-intensive workloads.
  • Optimize cluster performance, scheduling, and resource utilization using workload managers such as: Slurm, LSF, PBS Pro, Kubernetes

Security & Governance

  • Implement security best practices for HPC and cloud environments.
  • Manage IAM, secrets management, encryption, and compliance controls.
  • Support regulatory requirements and enterprise governance standards.

Qualifications

5+ years of experience in DevOps and Cloud infrastructure management

Technical Skills

  • Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or related field.
  • Strong experience with Linux system administration (RHEL, Rocky Linux, Ubuntu).
  • Experience managing HPC clusters and distributed computing environments.
  • Proficiency in Python, Bash, or Go.
  • Hands-on experience with: Terraform, Ansible, Git, Jenkins/GitHub Actions
  • Experience with container technologies: Docker, Kubernetes, Singularity/Apptainer
  • Knowledge of AI/ML frameworks: TensorFlow, PyTorch, Ray, Spark
  • Experience with GPU technologies and accelerator platforms.

Cloud Skills

  • AWS, Azure, or GCP architecture and operations.
  • Cloud networking, storage, and security services.
  • Hybrid cloud and HPC workload migration experience.

Additional Information

Sandisk thrives on the power and potential of diversity. As a global company, we believe the most effective way to embrace the diversity of our customers and communities is to mirror it from within. We believe the fusion of various perspectives results in the best outcomes for our employees, our company, our customers, and the world around us. We are committed to an inclusive environment where every individual can thrive through a sense of belonging, respect and contribution.

Sandisk is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at [email protected] to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.

Sandisk

About the company

Sandisk

Large Enterprise

SanDisk, a subsidiary of Western Digital Corporation, is a leading global provider of flash storage solutions. Founded in 1988, the company specializes in developing high-performance memory cards, USB flash drives, solid-state drives (SSDs), and enterprise storage solutions for devices ranging from smartphones to data centers. Known for its innovation and quality, SanDisk has been instrumental in advancing storage technology, enabling users to store and access vast amounts of data quickly and efficiently. With a commitment to reliability and cutting-edge design, SanDisk continues to empower individuals and businesses worldwide with its advanced storage products.