Mastercard

Mastercard

Posted via Workday

Principal Data Engineer (AWS, Databricks, Ai, ML Flow, Data Architecture, Apache Airflow)

Apply by Dec 31, 2026

Posted Sep 22, 2026

Role at a glance

Job function
AI & Data Data Engineering
Salary
Not Disclosed
Location
Pune, India
Work arrangement
On-site
Employment
Full-time
Experience
Typically 12-18 years of overall career relevant experience
Education
Bachelor's degree in Computer Science, Engineering, Information Systems, or a related technical discipline, or equivalent practical...

Spotted an issue?

We’ll check it against the original posting.

Log in to report

Role Summary

AI-generated

Mastercard Foundry R&D is seeking a Principal Data Engineer to design and build secure, scalable data foundations for advanced analytics, machine learning, generative AI, and agentic AI solutions. The role partners with AI Engineers, Data Scientists, Researchers, Software Engineers, and Product Managers to create reusable data capabilities and mature successful R&D initiatives into production solutions.

What You'll Do

  • Lead the architecture, design, and engineering of scalable batch, streaming, and event-driven data platforms.
  • Design and build data ingestion, transformation, enrichment, processing, and consumption pipelines.
  • Build and operate platforms supporting analytics, machine learning, generative AI, and agentic AI workloads.
  • Create reusable data products, services, APIs, and self-service capabilities.
  • Implement data governance capabilities including quality controls, lineage, metadata management, cataloging, access control, retention,...
  • Provide hands-on technical leadership through architecture reviews, code reviews, troubleshooting, and engineering mentorship.

Generated from the employer's posting. Verify important details before applying.

View full posting

Qualifications

Required: 12-18 years of relevant experience; large-scale production data platforms; Python, PySpark, SQL, Apache Spark, cloud-native data platforms, data architecture, batch and real-time pipelines, workflow orchestration, CI/CD, automated testing, version control, Infrastructure as Code, data governance, security, privacy, and technical leadership. Bachelor's degree in a related technical discipline or equivalent practical experience.

Required

  • 12-18 years of overall career relevant experience
  • Designing, building, and operating large-scale production data platforms
  • Expert programming skills in Python and PySpark
  • Advanced SQL expertise, including data modeling, performance tuning, query optimization, and large-scale analytical processing
  • Apache Spark and modern lakehouse platforms
  • Cloud-native data platforms in AWS, Azure, or other enterprise cloud environments
  • Modern data architecture patterns, including data lakes, lakehouses, data meshes, data products, and event-driven architectures
  • Cloud storage, data integration, streaming, serverless, observability, and security services

Preferred

  • Machine learning, generative AI, or agentic AI systems
  • MLflow or comparable ML lifecycle tooling
  • Azure Machine Learning, Azure AI Foundry, or Microsoft Fabric
  • Unity Catalog, Microsoft Purview, or comparable catalog and governance platforms
  • Kafka or other event-streaming technologies
  • Docker and Kubernetes
  • Terraform or another Infrastructure as Code framework
  • Structured, semi-structured, unstructured, and streaming data

Original job description

Content provided by the employer

Our Purpose

Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.

Title and Summary

Principal Data Engineer (AWS, Databricks, Ai, ML Flow, Data Architecture, Apache Airflow)

Overview
Mastercard Foundry R&D develops emerging technology solutions and transforms promising concepts into scalable products and platforms.
We are seeking a Principal Data Engineer to design and build the data foundations that power advanced analytics, machine learning, generative AI, and agentic AI solutions. This role requires a hands-on technical leader who can turn ambiguous R&D objectives into secure, scalable, and production-ready data systems.
You will collaborate with AI Engineers, Data Scientists, Researchers, Software Engineers, and Product Managers to create reusable data capabilities that accelerate experimentation and the delivery of innovative products.

Role
What You'll Do
• Lead the architecture, design, and engineering of scalable batch, streaming, and event-driven data platforms.
• Design and build reliable data ingestion, transformation, enrichment, processing, and consumption pipelines.
• Develop governed lakehouse and modern data platform architectures using cloud-native technologies and distributed data processing frameworks.
• Build and operate platforms that support analytics, machine learning, generative AI, and agentic AI workloads.
• Create reusable data products, services, APIs, and self-service capabilities that accelerate experimentation and production delivery.
• Enable the full AI and machine learning lifecycle through robust feature engineering, training, evaluation, deployment, and inference pipelines.
• Optimize data workloads for performance, scalability, reliability, resiliency, and cost efficiency.
• Define engineering standards for data modeling, software development, testing, CI/CD, observability, and operational excellence.
• Implement data governance capabilities including quality controls, lineage, metadata management, cataloging, access control, retention, and compliance.
• Apply security and privacy-by-design principles for sensitive and regulated data environments.
• Evaluate emerging data and AI technologies through prototypes, proof-of-concepts, and technical assessments.
• Translate loosely defined research, innovation, or product requirements into practical architectures and incremental delivery plans.
• Provide hands-on technical leadership through architecture reviews, code reviews, troubleshooting, and engineering mentorship.
• Partner with cross-functional teams to mature successful R&D initiatives into enterprise-grade production solutions.
• Communicate technical decisions, trade-offs, risks, dependencies, and roadmap recommendations to both technical and business stakeholders.
• Drive adoption of modern data engineering practices, platform automation, and platform reliability disciplines.
All About You
Required Qualifications
• Typically 12-18 years of overall career relevant experience, including significant ownership of complex enterprise data engineering solutions.
• Extensive experience designing, building, and operating large-scale production data platforms.
• Expert programming skills in Python, PySpark, and modern software engineering practices.
• Advanced SQL expertise, including data modeling, performance tuning, query optimization, and large-scale analytical processing.
• Strong hands-on experience with distributed data processing technologies such as Apache Spark and modern lakehouse platforms.
• Experience building and operating cloud-native data platforms in AWS, or Azure, or other enterprise cloud environments.
• Strong knowledge of modern data architecture patterns, including data lakes, lakehouses, data meshes, data products, and event-driven architectures.
• Experience with cloud storage, data integration, streaming, serverless, observability, and security services across public cloud platforms.
• Experience building scalable batch and real-time data pipelines.
• Experience with workflow orchestration platforms such as Apache Airflow, Databricks Workflows, AWS Step Functions, Azure Data Factory, or similar technologies.
• Experience implementing CI/CD, automated testing, version control, Infrastructure as Code, and platform automation.
• Strong understanding of data governance, metadata management, lineage, quality frameworks, privacy controls, and access management.
• Experience diagnosing and resolving complex performance, reliability, scalability, and operational challenges.
• Ability to make sound architectural decisions while balancing delivery speed, innovation, maintainability, security, and cost.
• Proven ability to thrive in R&D and innovation-focused environments where priorities and requirements may evolve through experimentation.
• Strong communication, collaboration, technical leadership, and mentoring skills.
• Bachelor's degree in Computer Science, Engineering, Information Systems, or a related technical discipline, or equivalent practical experience.

Preferred Qualifications
• Experience supporting machine learning, generative AI, or agentic AI systems.
• Experience with MLflow or comparable ML lifecycle tooling.
• Experience with Azure Machine Learning, Azure AI Foundry, or Microsoft Fabric.
• Experience with Unity Catalog, Microsoft Purview, or comparable catalog and governance platforms.
• Experience with Kafka or other event-streaming technologies.
• Experience with containerized and cloud-native platforms, including Docker and Kubernetes.
• Experience with Terraform or another Infrastructure as Code framework.
• Experience integrating structured, semi-structured, unstructured, and streaming data.
• Knowledge of responsible AI, model evaluation, and AI platform observability.
• Experience in payments, financial services, or another regulated industry.
Success in This Role
Success requires someone who:
• Remains strongly hands-on while providing technical direction.
• Can build production-quality systems without introducing unnecessary platform complexity.
• Converts experimentation into reusable engineering capabilities.
• Designs for security, governance, reliability, and observability from the outset.
• Challenges assumptions and validates architectural choices through evidence and prototypes.
• Enables AI and product teams rather than taking ownership of data science or model research.
• Influences across teams without depending on formal people-management authority.

Corporate Security Responsibility


All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:

  • Abide by Mastercard’s security policies and practices;

  • Ensure the confidentiality and integrity of the information being accessed;

  • Report any suspected information security violation or breach, and

  • Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.




Mastercard

About the company

Mastercard

Large Enterprise

Mastercard is a global technology company in the payments industry, committed to empowering individuals and businesses through secure and efficient payment solutions. With a presence in over 210 countries, Mastercard connects consumers, financial institutions, merchants, and governments, enabling seamless transactions across various platforms. The company is at the forefront of innovation, focusing on enhancing financial inclusion and expanding access to digital payment technologies. Through its advanced network and partnerships, Mastercard continues to revolutionize the way people engage in commerce worldwide.