Role at a glance
- Job function
-
AI & Data MLOps & ML Infrastructure
- Salary
- Not Disclosed
- Location
- Mississauga, Canada
- Employment
- Full-time
Spotted an issue?
We’ll check it against the original posting.
About the role
Original posting provided by AstraZeneca
WHY JOIN US?
Evinova is a health-tech business focused on accelerating better health outcomes by advancing digital transformation across the life sciences sector. By combining science-based expertise, evidence-led rigor, and deep human insight, we design digital solutions that enable healthcare to work better for everyone.
Operating at the intersection of healthcare, technology, data, and analytics, we are helping unlock the full potential of digital health, transforming how clinical research is conducted, how care is delivered, and how patients experience healthcare. Our solutions are built to scale, driving efficiency, improving decision-making, and ultimately delivering better outcomes for patients worldwide.
At Evinova, we are driven by a shared purpose to transform health through data and digital innovation. Our teams collaborate across disciplines to solve complex challenges, continuously learning and evolving in a fast-paced, high-impact environment.
We also recognize the importance of flexibility and balance. Our ways of working support both individual needs and team collaboration. To foster connection and collaboration, employees are expected to work from the office three days per week, creating opportunities for in-person teamwork, innovation, and meaningful connection.
This role is located in the Greater Toronto Area and follows a hybrid work model. Candidates must reside within commuting distance of the GTA or be willing to relocate for this opportunity.
Introduction to Role
The Machine Learning and Artificial Intelligence Operations team (ML/AI Ops) is the cloud platform engineering team responsible for building and operating the infrastructure that enables our AI Engineers and Data Scientists to deploy Generative AI applications reliably, securely, efficiently and at scale.
As a Senior Cloud Platform Engineer on the ML/AI Ops team, you will design, build and operate the AWS platform that runs our production Generative AI, agentic AI and conversational AI workloads. Most of the code you write will be AWS CDK (TypeScript and/or Python) that creates the infrastructure for other teams to build on.
This is a cloud and platform engineering role rather than an AI application development role. You will partner closely with the engineers and Data Scientists who build agents and models, and provide the infrastructure, deployment patterns, model access, observability and operational capabilities they need to move solutions from experimentation into reliable production environments.
You will work across AWS infrastructure, infrastructure as code (IaC), Amazon Bedrock AgentCore, Amazon ECS, CI/CD, SageMaker Unified Studio, AI gateways, observability, scalability, reliability, security, governance and cost optimization. Your work will establish reusable platform capabilities that allow teams across Evinova to deploy and operate solutions faster and more reliably while meeting the requirements of a highly regulated pharmaceutical environment.
Accountabilities
Cloud Platform Engineering
- Design, build and operate scalable AWS cloud platform capabilities for production ML/AI and Generative AI workloads.
- Create reusable infrastructure, tooling and deployment patterns that enable AI Engineers and Data Scientists to independently deploy and operate their applications.
- Write and maintain AWS IaC primarily AWS CDK in TypeScript and/or Python, including reusable CDK constructs that other teams consume.
- Build and operate containerized workloads using Amazon ECS and AgentCore.
- Develop reusable platform capabilities across compute, networking, IAM, secrets management, storage, model access and workload isolation.
- Build and maintain CI/CD and GitOps workflows that enable safe, automated and repeatable deployments across environments.
- Partner with engineers and Data Scientists to transition prototypes and research workloads into resilient, production-grade services.
- Build self-service capabilities and automation that improve developer experience and reduce operational toil.
Reliability, Scalability & Operational Excellence
- Engineer platform capabilities that improve the availability, scalability, resiliency and performance of production GenAI workloads.
- Design and implement autoscaling, load balancing, retries, timeouts, fallback, rate limiting and failure-recovery strategies.
- Establish monitoring, alerting, SLOs, runbooks and production-readiness standards for ML/AI workloads.
- Troubleshoot complex production issues across AWS infrastructure, container, application and model-provider layers.
- Automate operational processes and proactively identify opportunities to improve platform reliability and performance.
- Drive cloud and model cost optimization through capacity management, workload optimization and data-driven analysis.
AI Gateway & Model Access
- Build and operate a centralized AI gateway that routes requests across model providers.
- Provide secure, reliable and governed access to foundation models through Amazon Bedrock, OpenAI, Anthropic, Microsoft Foundry, Gemini Enterprise Agent Platform (formerly Vertex AI) and other model platforms.
- Implement model/provider routing, fallback, authentication, rate limiting, quotas and cost controls.
- Enable AI teams to evaluate and change model providers without tightly coupling their applications to individual model endpoints.
- Provide the compute, networking, storage and runtime infrastructure to reliably operate agentic AI, RAG and conversational AI workloads in production.
Observability, Governance & Cost Optimization
- Build platform-level observability for GenAI workloads, including token consumption, latency, throughput, errors, model/provider performance and cost attribution by team and application.
- Implement standardized tracing, logging, metrics and alerting capabilities that can be adopted across AI applications.
- Integrate observability technologies such as Amazon CloudWatch, OpenTelemetry, Datadog and Splunk.
- Build the telemetry and data pipelines that AI teams use to evaluate and monitor production LLM behavior.
- Build appropriate security, auditability and governance controls for AI workloads operating within a regulated environment.
- Support compliance with applicable industry standards and practices, including Good Clinical Practice and Good Machine Learning Practice and other GxP related standards and practices.
Representative Projects
- Build a library of AWS CDK constructs that gives an AI team a production-ready Amazon ECS service, with IAM, secrets, networking and observability, from a single import.
- Deploy and operate a centralized AI gateway on Amazon ECS, with multi-provider routing, fallback, quotas and rate limiting.
- Build multi-account CI/CD with CDK Pipelines or GitOps, with safe, staged deployments across environments.
- Implement token cost attribution by team and application with SLO and burn-rate alerts.
- Migrate container workloads from Amazon EKS to Amazon ECS without loss of reliability.
Essential Skills/Experience
- Minimum of 4 years of hands-on experience in a platform engineering, infrastructure, site reliability engineering (SRE) or DevOps role where infrastructure code was your main output.
- Deep hands-on AWS cloud engineering experience, including designing, deploying and operating production cloud-native infrastructure.
- Strong experience with AWS services such as Bedrock, ECS, IAM, VPC, load balancing, S3, CloudWatch, Secrets Manager and related services.
- Strong hands-on experience with IaC. AWS CDK using Python and/or TypeScript is preferred. Strong Terraform engineers who are willing to work in CDK are welcome.
- Strong experience with Docker and Amazon ECS, including workload deployment, autoscaling, resource management, health checks, networking, security and production troubleshooting.
- Experience designing and maintaining CI/CD and/or GitOps pipelines for production cloud workloads.
- Strong understanding of cloud and platform engineering principles, including high availability, scalability, resiliency, observability, performance, security and cost optimization.
- Experience with on-call rotations, incident response and root-cause analysis for production systems.
- Strong software engineering skills in Python and/or TypeScript, with experience building production-quality platform tooling and automation.
- Experience implementing production observability using technologies such as CloudWatch, OpenTelemetry, Prometheus, Splunk, Datadog, Pydantic Logfire, Langfuse or comparable tools.
- Demonstrated experience partnering with software, data or ML/AI teams to transition workloads from experimentation into reliable production environments.
- Proven ability to collaborate effectively across ML/AI engineering, software engineering, security, infrastructure and product teams.
- Strong written and verbal communication skills, including the ability to clearly document infrastructure architecture, operational processes and platform standards.
Preferred Skills/Experience
- Experience building internal developer platforms or self-service cloud capabilities used by multiple engineering teams.
- Experience supporting Generative AI/LLM workloads in production, with an understanding of their unique operational challenges including model availability, latency, token consumption, evaluation and cost.
- Hands-on experience with Amazon Bedrock or another enterprise foundation-model platform.
- Experience implementing and/or configuring AI/model gateways, proxies or routing layers, including model routing, fallback, rate limiting and provider abstraction.
- Experience operating infrastructure supporting agentic AI, multi-agent systems, RAG pipelines or conversational AI.
- Experience with AgentCore or comparable managed agent-runtime capabilities.
- Experience supporting Model Context Protocol (MCP) tools, servers or services.
- Experience with LLM evaluation and observability platforms such as Arize Phoenix, Langfuse, Braintrust, Freeplay or Logfire.
- Experience supporting production ML/AI infrastructure within pharmaceutical, healthcare, financial services or another highly regulated industry.
- Understanding of GxP and/or other controls applicable to regulated ML/AI systems.
- Experience with Kubernetes (Amazon EKS), including migrating container workloads to Amazon ECS.
This Role Is Not a Good Fit If
- Your main experience is building agents, RAG pipelines, prompts or LLM applications, and you have not owned production infrastructure.
- You have not written or maintained infrastructure as code for production systems.
- You prefer research or model development to operating production platforms.
Personal Attributes
- Platform-minded: You think beyond individual applications and build reusable capabilities that enable multiple engineering teams.
- Operationally focused: You care about reliability, scalability, observability, security, performance and what happens after a workload reaches production.
- Hands-on: You are comfortable writing code, building infrastructure, operating container platforms and troubleshooting production systems.
- Customer-obsessed: You view the AI Engineers and Data Scientists using the platform as your customers and continually look for ways to improve their developer experience.
- Organized and attentive to detail: You effectively manage multiple initiatives and competing priorities.
- Collaborative and inclusive: You foster a positive engineering culture where teams share knowledge and solve problems together.
- Comfortable working autonomously: You like working in an evolving environment where you can help establish new systems, patterns, and best practices.
- Curious and committed to staying current: You love learning and applying that knowledge to the systems you build. You stay current with the latest AI trends and technologies and apply that knowledge to your work pragmatically.
SO, WHAT’S NEXT?
To be considered for this exciting opportunity, please complete the full application on our website at your earliest convenience – it is the only way that our Recruiter and Hiring Manager can know that you feel well qualified for this opportunity. If you know someone who would be a great fit, please share this posting with them.
Where can I find out more?
- Explore what we’re building: www.evinova.com
- Stay connected and see our impact in action: https://www.linkedin.com/company/evinova/
Apply today to bring smarter, faster clinical trials to life!
Evinova is an equal opportunity employer that is committed to diversity and inclusion and providing a workplace that is free from discrimination. Evinova is committed to accommodating persons with disabilities. Such accommodation is available on request in respect of all aspects of the recruitment, assessment and selection process and may be requested by emailing [email protected].
#LI-Hybrid
⠀
Annual base salary for this position ranges from 134,855.20 to 176,997.45.AstraZeneca is committed to providing fair and equitable compensation opportunities to all colleagues. Our compensation policies and practices have been designed to allow colleagues to progress through the salary range over time as they progress in their role. The range provided in this posting represents an offer pay range used in a majority of situations. The base pay offered will vary depending on multiple individualized factors, including the candidate's skills and experience, job-related knowledge, and other specific business and organizational needs. In some cases, offers outside the range may also be considered to address unique circumstances.
In addition, our permanent positions offer an annual Variable Pay Bonus/Short Term Incentive opportunity as well as eligibility to participate in our equity-based long-term incentive program (if applicable to role). Benefits offered for permanent roles include a competitive Flex Benefits & Retirement Savings Program, 4 weeks’ paid vacation, and annual Personal Days. Fixed Term Contract/Temporary positions (excluding students) are offered a Contract Benefits Program.
We are using AI as part of the recruitment process.
This advertisement relates to a current vacancy.
About the company
AstraZeneca
Large Enterprise
AstraZeneca is a global biopharmaceutical company dedicated to the research, development, and commercialization of innovative medicines. Headquartered in Cambridge, UK, the company focuses on areas such as oncology, cardiovascular, renal, metabolism, and respiratory diseases, striving to deliver life-changing treatments to patients. With a strong commitment to scientific excellence and collaboration, AstraZeneca invests in cutting-edge technologies and partnerships to drive advancements in healthcare. The company is also committed to sustainability and corporate responsibility, aiming to make a positive impact on communities worldwide.