Microsoft

Microsoft

Posted via Microsoft Careers

Member of Technical Staff - AI Evaluations, Health

Posted Aug 26, 2026

Role at a glance

Salary
$119.8K – $234.7K/yr
Location
New York, New York, United States NY, New York, United States
Work arrangement
Hybrid
Employment
Full-time
Education
Bachelor's

Spotted an issue?

We’ll check it against the original posting.

Log in to report

About the role

Original posting provided by Microsoft

View original
Overview

MAI Health is bringing world-class AI to healthcare, and this role puts you at the front line of that effort: embedded in our partnership with Mayo Clinic, one of the world's leading medical institutions. As an Evals Engineer, you will work with Mayo Clinic to ensure the agents we deploy are safe and effective. These agents help patients with everything from figuring out whether Mayo is the right fit for their care, to navigating their visit, understanding their conditions, and answering billing questions. This role is critical to defining what "good" looks like for a Mayo Clinic agent starting from an explicit specification of intended behavior, grounded in good medical practice and building the evaluation system that ensures quality improves over time.

Being passionate and opinionated about human-computer interaction, you will work at the nexus of product, research, and clinical practice. Your job is to evaluate the configured system patients actually encounter: the model together with its orchestration harness, tools, and connected data. You will be responsible for ensuring outputs are high quality, factual, and safe in a domain where the bar for all three is the highest anywhere. We're looking for someone with an abundance of positive energy, empathy, and kindness, in addition to being highly effective. The right candidate takes initiative, thrives with ambiguity, and is comfortable being the face of MAI inside a partner organization.

This is a hybrid role based in NYC, with regular travel to Mayo Clinic (Rochester, MN) for on-site working sessions with clinical and technical stakeholders, and to London to work with the larger MAI Health Evals team.



Responsibilities
  • Own the end-to-end evaluation strategy for MAI Health's Mayo Clinic project: defining metrics, building datasets, and translating results into actionable steps for response-quality improvement, with clinical accuracy and patient safety as the top priorities.
  • Build layered evaluations that span the full interaction spectrum from static single-turn benchmarks, adaptive multi-turn simulations with user simulators (including adversarial cases with hidden goals), and grounded, agentic evaluations that exercise the agent's tools, retrieval, and connected data, scoring trajectories as well as final responses.
  • Be a strong member of the MAI Health Evaluations team: learn from its practices and frameworks, translate them to the Mayo context, and feed what you learn back to improve the overall team's capabilities.
  • Work embedded with Mayo clinicians and subject-matter experts to design high-quality tooling and run evaluations for LLM-based tools. This includes writing clinically-grounded evaluation guidelines and exemplar-tied, case-specific rubrics, recruiting and calibrating physician raters, and maintaining held-out clinician-authored test sets to detect overfitting.
  • Design and write LLM-as-judge prompts for medical content; build and run automated evaluations that scale clinical review, calibrate autograders against expert clinician annotation, and report judge–human agreement alongside inter-rater agreement so the measurement itself is trustworthy.
  • Evaluate patient-facing AI experiences for factuality, groundedness, safety, appropriate escalation, and health literacy. This includes safeguards for psychologically vulnerable users and safety-netting behaviors and ensuring outputs are trustworthy whether the customer is a consumer or a clinician.
  • Serve as the technical partner ensuring high-quality evaluations are created at Mayo: gathering requirements, demonstrating capabilities, triaging quality issues, and representing partner needs back to MAI engineering and research teams.
  • Partner with data teams to develop scalable pipelines and dashboards tracking pre-launch and in-deployment eval performance.
  • Contribute to privacy-preserving monitoring of production conversations, pairing it with clinical review of concerning interactions, so evaluation evolves with real-world use.
  • Translate evaluation insights into actionable product improvements and inform launch/no-launch decisions for clinical features, maintaining configuration-specific records (model version, harness, evaluation date) that support lifecycle oversight and change control.
  • Monitor trends in model performance, clinician and patient feedback, and the clinical-AI benchmark landscape (e.g., medical QA benchmarks, safety frameworks) to continuously refine evaluation approaches.
  • Own key projects end-to-end, proactively identifying risks and proposing solutions to ensure timely delivery.
  • Coordinate cross-team and cross-organization collaboration: align MAI engineering, research, data science, and Mayo clinical/IT stakeholders on goals, deliverables, and status.
  • Handle sensitive health data responsibly, working within HIPAA and partner data-governance requirements.
  • Embody our Culture and Values.


Qualifications

Required Qualifications

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience
  • 3+ years leading multi-disciplinary projects: defining requirements, developing project plans, and executing with cross-functional teams.

    Preferred Qualifications: 

  • Experience in healthcare, health tech, or other regulated, safety-critical domains; familiarity with clinical workflows, medical terminology, or working with clinician stakeholders.
  • Comfort operating inside a partner organization and translating between their world and yours.
  • Computer science background; experience in training/evaluation of LLMs, including LLM-as-judge design, user simulation, and agentic/trajectory evaluation.
  • Technical depth in software development, data science, and machine learning. While you are not expected to write code on the critical path, you're able to define and execute zero-to-one prototypes, build tools to support the larger eval pipeline, and speak the language of the engineers, researchers, and clinicians you work with.
  • Familiarity with HIPAA, PHI handling, and/or lifecycle regulatory approaches to clinical AI (e.g., FDA guidance on AI-enabled devices and predetermined change-control plans, EU AI Act).
  • Experience building and evaluating ML-powered or LLM-powered products.


Software Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay

Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Microsoft

About the company

Microsoft

Large Enterprise

Microsoft is a global technology leader that empowers individuals and organizations to achieve more through innovative software, services, and devices. Founded in 1975, the company is best known for its flagship products like the Windows operating system and Microsoft Office suite. In addition to personal computing, Microsoft is a leader in cloud computing with its Azure platform, providing a range of solutions for businesses to enhance productivity and efficiency. With a strong commitment to sustainability and accessibility, Microsoft continues to drive technological advancements that shape the future of work and learning.