Role at a glance
- Job function
-
AI & Data Data Engineering
- Salary
- $154.6K – $209.1K/yr
- Location
- Boulder, Colorado, United States
- Work arrangement
- On-site
- Employment
- Full-time
Spotted an issue?
We’ll check it against the original posting.
Qualifications
Required
- 5+ years of data engineering experience
- Experience with data modeling, warehousing and building ETL pipelines
- Experience with SQL
- Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
- Experience mentoring team members on best practices
Preferred
- Experience with big data technologies such as: Hadoop, Hive, Spark, EMR
- Experience operating large data warehouses
About the role
Original posting provided by amazon
Are you excited by the idea of building the data foundation that an entire AI-powered product is built on — from the very first pipeline? Do you like the messy, ambiguous problems: wrangling inconsistent data from dozens of outside partners and turning it into something clean, timely, and genuinely trustworthy? If so, we'd love to talk.
Marketing at Amazon happens across a huge range of independent businesses, and the data behind it is scattered across agencies, vendors, and platforms. It's slow to get, inconsistent, and almost impossible to see as a whole. We're a new, AI-native team setting out to fix that — building a service that automatically pulls all of this data together, cleans and normalizes it, and puts AI-powered analytics on top so anyone can ask a question in plain language and get a real answer.
That's where you come in. As a Data Engineer on this team, you'll own the ingestion pipelines and the normalized dataset that everything else depends on. You'll figure out how to reliably pull data from noisy, ever-changing external sources, reconcile feeds that never quite agree with each other, and build the quality checks that catch problems before anyone downstream ever sees them. AI is only as good as the data underneath it, so your work sits squarely on the critical path — if the data foundation is solid, the whole product wins.
You won't just be maintaining pipelines. You'll help shape how data flows through a brand-new system and decide what "good" looks like for schemas and data quality. It's a small, scrappy team with a lot of ownership to go around, and you'll get to leave your fingerprint on something from the ground floor.
If building trustworthy data infrastructure for hard, never-been-done-before problems sounds like your kind of challenge, we'd love to hear from you.
Key job responsibilities
- Design, build, and operate scalable, reliable data ingestion pipelines that collect marketing data from agencies, ad tech vendors, and internal sources on scheduled intervals — without requiring source teams to change their current processes.
- Define and evolve input schemas and normalization logic that reconcile inconsistent granularity, formats, and terminology across sources into a unified dataset.
- Build automated data quality validation, anomaly detection, and reconciliation to catch and surface defects early, reducing manual auditing and cleaning effort.
- Handle real-world data complexity, such as delayed settlement of certain data sources, mapping records to the correct hierarchies, and deciding when to use planned versus finalized data.
- Partner with modeling and AI engineers to ensure the data foundation is structured for downstream analytics, agent orchestration, and natural-language query.
- Instrument pipelines for observability, freshness/SLA monitoring, and low-touch operations, so the service requires primarily configuration updates as new sources are added.
- Collaborate with data providers and stakeholders to onboard new feeds and improve their timeliness, completeness, and accuracy.
- Contribute to a phased rollout, starting with a core set of sources and expanding over time.
A day in the life
Your morning might start by reviewing pipeline health and resolving a data quality alert before anyone downstream is affected. Mid-morning, you onboard a new external source — reverse-engineering a messy partner feed and designing the schema to fold it cleanly into the shared dataset. After lunch, you meet with a data provider to close gaps in their feed, then review a teammate's pull request. You close the day prototyping a better way to catch bad records at ingestion. Because this service is being built from scratch, you operate with real autonomy and set the patterns others build on.
Basic Qualifications
- 5+ years of data engineering experience- Experience with data modeling, warehousing and building ETL pipelines
- Experience with SQL
- Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
- Experience mentoring team members on best practices
Preferred Qualifications
- Experience with big data technologies such as: Hadoop, Hive, Spark, EMR- Experience operating large data warehouses
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, CO, Boulder - 154,600.00 - 209,100.00 USD annually
About the company
amazon
Large Enterprise
Amazon is a global leader in e-commerce and cloud computing, founded in 1994 by Jeff Bezos. Initially starting as an online bookstore, it has since expanded its offerings to include a vast range of products and services, including electronics, fashion, and digital content. With Amazon Web Services (AWS), the company also provides powerful cloud solutions to businesses around the world. Known for its innovation, customer-centric approach, and commitment to operational efficiency, Amazon continues to shape the future of retail and technology, consistently seeking new ways to enhance customer experiences.