Role at a glance
- Job function
-
AI & Data Data Engineering
- Salary
- Not Disclosed
- Location
- TS, Indiana, United States
- Employment
- Full-time
Spotted an issue?
We’ll check it against the original posting.
Role Summary
The Data Engineer develops and maintains data pipelines and infrastructure that support the organization’s data needs. The role focuses on reliable data ingestion, performance, quality, and security, while working with data and platform teams.
What You'll Do
- Develop and maintain ETL/ELT pipelines to ingest data from various sources into the data warehouse.
- Optimize data storage and retrieval for performance and scalability.
- Collaborate with data architects, analysts, and scientists to understand data needs and ensure the infrastructure supports them.
- Ensure data quality and integrity through validation and testing.
- Implement and maintain security protocols to protect sensitive data.
- Partner with the Enterprise Data and Analytics Platform team, functional data teams, and the Data Community lead to enable adoption of...
Generated from the employer's posting. Verify important details before applying.
View full postingQualifications
Requires 5+ years of experience in data engineering or software development, including hands-on experience implementing and operating data capabilities and solutions. Required expertise includes Databricks; designing real-time data ingestion pipelines in AWS; CloudFormation and custom resources in CFTs; GitHub Workflows integrated with AWS; building APIs using AWS services; and AWS Glue and its data engineering ecosystem. Also seeks programming skills in languages such as Python, R, PyTorch, PySpark, Pandas, or Scala; SQL and database technologies; experience with AWS, Azure, or Google Cloud Platform; and experience in an Agile/product-based environment. Strong analytical, problem-solving, communication, and collaboration skills are also listed.
Required
- Hands-on experience implementing and operating data capabilities and solutions; 5+ years of experience in data engineering or software...
- Expertise in Databricks.
- Expertise designing and building real-time data ingestion pipelines in AWS.
- Expertise in CloudFormation and developing and using custom resources in CFTs.
- Expertise developing GitHub Workflows and integrating them with AWS.
- Expertise building APIs using AWS services, and in AWS Glue and the AWS data engineering ecosystem.
- Strong programming skills in languages such as Python, R, PyTorch, PySpark, Pandas, or Scala.
- Experience with SQL and database technologies such as MySQL, PostgreSQL, or Presto.
Preferred
- Cloud environment experience is preferred.
- Boto3 API experience for Lambda, S3, Glue, and Crawlers; Data Zone is specifically marked preferred.
- Hands-on experience delivering data and ETL solutions using AWS data services such as Redshift, Athena, and Lake Formation, Cloudera...
- Functional knowledge or prior experience in the Life Sciences Research and Development domain is a plus.
Original job description
Content provided by the employer
Original job description
Content provided by the employer
At Bristol Myers Squibb, our employees often ask, “Who are you working for?”—a question that fuels collaboration, accountability, and urgency in our work. Our purpose-driven culture inspires us to discover, develop, and deliver innovative medicines to prevail over serious diseases. We offer uniquely interesting and meaningful work, opportunities for growth, and a supportive environment that values inclusion, wellbeing, flexibility, and comprehensive benefits. This is work that transforms the lives of patients, and the careers of those who do it.
Key Responsibilities
The Data Engineer will be responsible for developing and maintaining ETL/ELT pipelines for ingesting data from various sources into our data warehouse.
Work with an end-to-end ownership mindset, innovate and drive initiatives through completion.
Optimize data storage and retrieval to ensure efficient performance and scalability
Collaborate with data architects, data analysts and data scientists to understand their data needs and ensure that the data infrastructure supports their requirements
Ensure data quality and integrity through data validation and testing
Implement and maintain security protocols to protect sensitive data
Stay up-to-date with emerging trends and technologies in data engineering and analytics
Closely partner with the Enterprise Data and Analytics Platform team, other functional data teams and Data Community lead to enable adoption of data and technology strategy.
Knowledgeable in evolving trends in Data platforms and Product based implementation
Comfortable working in a fast-paced environment with minimal oversight
Prior experience working in an Agile/Product based environment.
Qualifications & Experience
3-5 years of hands-on experience working on implementing and operating data capabilities and cutting-edge data solutions, preferably in a cloud environment. Breadth of experience in technology capabilities that span the full life cycle of data management including data lakehouses, master/reference data management, data quality and analytics/AI ML is needed.
Expertise in Databricks
Expertise in designing and building real time data ingestion data pipelines in AWS
Expertise in CloudFormation
Developing and using Custom Resources in CFTs
Expertise in developing GitHub Workflows & integrating GitHub workflows with AWS
Expertise with using boto3 apis for Lambda, S3, Glue, Crawlers, Data Zone (preferred)
Expert in building APIs using AWS services
In-depth knowledge and hands-on experience with ASW Glue services and AWS Data engineering ecosystem.
Hands-on experience developing and delivering data, ETL solutions with some of the technologies like AWS data services (Redshift, Athena, lakeformation, etc.), Cloudera Data Platform, Tableau labs is a plus
5+ years of experience in data engineering or software development
Create and maintain optimal data pipeline architecture, assemble large, complex data sets that meet functional / non-functional business requirements.
Identify, design, and implement internal process improvements: automating manual processes, optimizing data delivery, re-designing infrastructure for greater scalability, etc.
Strong programming skills in languages such as Python, R, PyTorch, PySpark, Pandas, Scala etc.
Experience with SQL and database technologies such as MySQL, PostgreSQL, Presto, etc.
Experience with cloud-based data technologies such as AWS, Azure, or Google Cloud Platform
Strong analytical and problem-solving skills
Excellent communication and collaboration skills Functional knowledge or prior experience in Lifesciences Research and Development domain is a plus
Experience and expertise in establishing agile and product-oriented teams that work effectively with teams in US and other global BMS site.
Initiates challenging opportunities that build strong capabilities for self and team
Demonstrates a focus on improving processes, structures, and knowledge within the team. Leads in analyzing current states, deliver strong recommendations in understanding complexity in the environment, and the ability to execute to bring complex solutions to completion.
We hire for skills and capabilities, not just credentials – if this role excites you, but doesn’t perfectly match your resume, we encourage you to apply anyway.
How We Work
Where you work matters – because collaboration, innovation and patient impact happen in many settings. Our roles are structured across four work models: site-essential, site-by-design, field-based and remote-by-design. The model assigned to this role is based on its core responsibilities. Learn more at https://careers.bms.com/ways-of-working.
Supporting People with Disabilities
BMS is dedicated to ensuring that people with disabilities can excel through a transparent recruitment process, reasonable workplace accommodations/adjustments and ongoing support in their roles. Applicants can request a reasonable workplace accommodation/adjustment prior to accepting a job offer. If you require reasonable accommodations/adjustments in completing this application, or in any part of the recruitment process, direct your inquiries to [email protected]. Visit careers.bms.com/eeo-accessibility to access our complete Equal Employment Opportunity statement.
Candidate Rights
BMS will consider qualified applicants with arrest and conviction records, pursuant to applicable laws in your area.
For roles based in Los Angeles County only: If you live in or expect to work from Los Angeles County if hired for this position, please visit this page for important additional information: https://careers.bms.com/california-residents/
Data Protection
We will never request payments, financial information, or social security numbers during our application or recruitment process. Learn more about protecting yourself at https://careers.bms.com/fraud-protection.
Any data processed in connection with role applications will be treated in accordance with applicable data privacy policies and regulations.
If this posting is missing required information required by local law or incorrect, contact BMS at [email protected] with the Job Title and Requisition number. Do not send application-related inquiries to this email. To check your application status, please login to your Candidate Home Account.
About the company
Bristol Myers Squibb
Large Enterprise
Bristol Myers Squibb is a global biopharmaceutical company committed to discovering, developing, and delivering innovative medicines that help patients prevail over serious diseases. With a focus on oncology, immunology, and cardiovascular health, the company leverages advanced science and technology to create transformative therapies that address critical unmet medical needs. Operating in more than 70 countries, Bristol Myers Squibb emphasizes collaborative efforts and strong partnerships in research and development to drive advancements in healthcare and improve patient outcomes worldwide.