Role at a glance
- Job function
-
AI & Data MLOps & ML Infrastructure AI Solutions Architecture
- Salary
- Not Disclosed
- Location
- Bengaluru
- Work arrangement
- On-site
- Employment
- Full-time
- Experience
- Minimum 7.5 year(s) of experience is required
- Education
- 15 years full time education
Spotted an issue?
We’ll check it against the original posting.
Role Summary
This role is a Senior Engineer in AI Infrastructure Architecture for AWS, responsible for designing and engineering optimized compute infrastructure for large-scale AI and machine learning systems. The work includes building scalable training and model-serving foundations, automation patterns and operational controls aligned with client standards, SLAs, security, compliance and cost-efficiency expectations.
What You'll Do
- Own the end-to-end architecture and design of AWS compute infrastructure for large-scale AI/ML systems, including distributed training,...
- Design and tune AWS GPU clusters and distributed training systems using services including EC2, EKS, SageMaker, S3, FSx/EFS, VPC, IAM...
- Evaluate architecture alternatives and lead reviews of existing and proposed environments, identifying risks, bottlenecks and...
- Define infrastructure roadmaps, capacity planning models, scaling strategies, cost forecasts and performance improvement opportunities.
- Design deployment, automation and CI/CD strategies for reliable releases of AI systems, models, data pipelines and platform components.
- Establish monitoring and observability practices across InfraOps and MLOps, including SLAs, SLOs, alerting and performance and cost...
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Required qualifications include a bachelor's degree in Computer Science, Computer Engineering, Information Technology or a related engineering field; AWS AI infrastructure experience; AI/ML infrastructure and cloud platform experience; programming or scripting proficiency; data pipeline and workflow management experience; and experience with distributed training, GPU-accelerated compute, model serving, containers, networking, CI/CD, Terraform or CloudFormation, observability, and incident response.
Required
- AWS AI Services
- Bachelor's degree in Computer Science, Computer Engineering, Information Technology or a related engineering field
- Minimum 4 years of experience coding, building, monitoring, troubleshooting, designing and operating AI/ML infrastructure, cloud...
- Strong understanding of AI/ML concepts and computing infrastructure for production AI workloads
- Minimum 4 years of proficiency in programming or scripting languages such as Python, Java, C++, Bash or PowerShell
- Experience with Apache Airflow, Kubeflow, managed orchestration services or platform-native workflow tooling
- Minimum 4 years of experience in AI/ML infrastructure engineering or related roles on a hyperscaler or enterprise platform
- Experience with AWS AI infrastructure services including EC2, EKS, SageMaker, S3, FSx/EFS, IAM, VPC, CloudWatch and AWS DevOps/security...
Preferred
- Machine Learning Operations
- AWS certifications such as Solutions Architect Professional, DevOps Engineer Professional or Machine Learning/AI specialty or associate...
- Industry experience in BFSI, healthcare, retail/e-commerce, telecom, manufacturing, energy or public sector environments
- Exposure to LLM infrastructure, vector databases, retrieval pipelines, GPU scheduling, high-performance storage, low-latency model...
- Knowledge of enterprise architecture governance, FinOps, infrastructure partner/vendor collaboration and production support operating models
Original job description
Content provided by the employer
Original job description
Content provided by the employer
Project Role Description : Architects the data platform blueprint and implements the design, encompassing the relevant data platform components. Collaborates with the Integration Architects and Data Architects to ensure cohesive integration between systems and data models.
Must have skills : AWS AI Services
Good to have skills : Machine Learning Operations
Minimum 7.5 year(s) of experience is required
Educational Qualification : 15 years full time education
Role Summary / Description
AI Powered Tech Talent
As a Senior Engineer in AI Infrastructure Architecture for AWS, you will own significant portions of the end-to-end architecture and engineering of optimized compute infrastructure for large-scale AI and machine learning systems. You will design scalable distributed training environments, model-serving foundations, automation patterns and operational controls that align with client standards, SLAs, security, compliance and cost-efficiency expectations. You will bring industry experience across enterprise AI adoption, cloud modernization, regulated workloads, FinOps and production reliability, while mentoring engineers and partnering with architects to translate business requirements into robust AWS-based AI infrastructure solutions.
Key Responsibilities
Own end-to-end architecture and design of optimized AWS compute infrastructure for large-scale AI/ML systems, including distributed training, GPU/accelerated compute, container platforms and model-serving environments.
Design and tune large-scale AWS GPU clusters and distributed training systems using services such as EC2, EKS, SageMaker, S3, FSx/EFS, VPC, IAM and CloudWatch, including accelerator selection, interconnect/networking and high-throughput storage design.
Serve as an authoritative AI infrastructure expert on AWS, applying deep knowledge of AWS AI/ML services, accelerators, networking, security and cost levers.
Develop and evaluate architecture alternatives, weighing trade-offs across compute, networking, storage, orchestration, model serving, observability, security, compliance, cost and operational complexity.
Lead architecture assessments and reviews of existing and proposed environments, identifying gaps, risks, bottlenecks and optimization opportunities, and recommending remediation actions.
Drive architecture decision-making by documenting rationale, trade-offs, assumptions and dependencies so decisions are transparent, defensible and aligned with business SLAs and standards.
Define and maintain AI infrastructure roadmap inputs, capacity planning models, scaling strategies, cost forecasts and performance improvement opportunities.
Design deployment, automation and CI/CD strategies for reliable, repeatable and scalable releases of AI systems, models, data pipelines and platform components into production.
Establish AI monitoring and observability practices across InfraOps and MLOps, including SLAs, SLOs, alerting, performance/cost tracking and continuous optimization.
Integrate AI/ML systems into enterprise environments while ensuring interoperability, security, compliance, regulatory alignment and adherence to client standards.
Collaborate with clients, stakeholders, architects and engineering teams to align infrastructure decisions with business outcomes and translate requirements into actionable architecture standards.
Set technical direction for workstreams, mentor engineers, review designs/code and promote engineering best practices across the team.
Required Qualifications
Bachelor's degree in Computer Science, Computer Engineering, Information Technology or a related engineering field.
Minimum 4 years of experience coding, building, monitoring, troubleshooting, designing and operating AI/ML infrastructure, cloud platforms, data platforms, model deployment pipelines or large-scale engineering solutions.
Strong understanding of AI/ML concepts and the computing infrastructure required to deploy, run and optimize production AI workloads.
Minimum 4 years of proficiency in programming or scripting languages such as Python, Java, C++, Bash, PowerShell or equivalent engineering languages.
Experience with data pipeline and workflow management tools such as Apache Airflow, Kubeflow, managed orchestration services or platform-native workflow tooling.
Strong problem-solving skills and ability to work in a fast-paced engineering or client delivery environment.
Excellent communication, collaboration and stakeholder alignment skills.
Minimum 4 years of experience in AI/ML infrastructure engineering or related roles on a hyperscaler or enterprise platform for deploying large-scale solutions.
Proven experience leading AI projects or engineering workstreams and managing priorities across multiple initiatives.
Demonstrated experience evaluating and selecting AI technologies, frameworks, cloud services and architecture patterns.
Required Skills/ Experience
Strong hands-on experience with AWS AI infrastructure services including EC2, EKS, SageMaker, S3, FSx/EFS, IAM, VPC, CloudWatch and AWS DevOps/security services.
Experience architecting GPU/accelerated compute, distributed training, model serving, high-throughput storage, container platforms and secure cloud networking.
Strong working knowledge of Terraform/CloudFormation, CI/CD, Docker, Kubernetes, InfraOps, MLOps, observability and incident response practices.
Ability to optimize AWS AI infrastructure for performance, power, cost, scalability, security, reliability and compliance.
Experience producing architecture decision records, reference implementations, standards, runbooks and reusable infrastructure patterns.
Good to Have Skills
AWS certifications such as Solutions Architect Professional, DevOps Engineer Professional or Machine Learning/AI specialty or associate credentials.
Industry experience in BFSI, healthcare, retail/e-commerce, telecom, manufacturing, energy or public sector environments where AI infrastructure must meet compliance, security, reliability and cost-control requirements.
Exposure to LLM infrastructure, vector databases, retrieval pipelines, GPU scheduling, high-performance storage, low-latency model serving and model optimization techniques.
Knowledge of enterprise architecture governance, FinOps, infrastructure partner/vendor collaboration and production support operating models.15 years full time education
About Accenture
Accenture is a leading global professional services company that helps the world’s leading businesses, governments and other organizations build their digital core, optimize their operations, accelerate revenue growth and enhance citizen services—creating tangible value at speed and scale. We are a talent- and innovation-led company with approximately 791,000 people serving clients in more than 120 countries. Technology is at the core of change today, and we are one of the world’s leaders in helping drive that change, with strong ecosystem relationships. We combine our strength in technology and leadership in cloud, data and AI with unmatched industry experience, functional expertise and global delivery capability. Our broad range of services, solutions and assets across Strategy & Consulting, Technology, Operations, Industry X and Song, together with our culture of shared success and commitment to creating 360° value, enable us to help our clients reinvent and build trusted, lasting relationships. We measure our success by the 360° value we create for our clients, each other, our shareholders, partners and communities.Visit us at www.accenture.com
Equal Employment Opportunity Statement
We believe that no one should be discriminated against because of their differences. All employment decisions shall be made without regard to age, race, creed, color, religion, sex, national origin, ancestry, disability status, military veteran status, sexual orientation, gender identity or expression, genetic information, marital status, citizenship status or any other basis as protected by applicable law. Our rich diversity makes us more innovative, more competitive, and more creative, which helps us better serve our clients and our communities.
About the company
Accenture
Large Enterprise
Accenture is a global professional services company that specializes in providing consulting, technology, and outsourcing services. With a diverse range of industries served, including financial services, healthcare, and telecommunications, Accenture leverages advanced technologies and data analytics to help organizations improve their performance and drive innovation. Committed to sustainable progress, the company emphasizes its dedication to inclusivity, digital transformation, and building a more sustainable future for its clients and communities. With a presence in over 120 countries, Accenture is known for its expertise in integrating cutting-edge solutions that address complex business challenges.