Multiple Locations, United States

Hardware Health

Location
Multiple Locations, United States
Job Number
200009249-en-1
City
Multiple Locations
Team
MAI AI Platform
Country
United States
Discipline
Software Engineering
Overview

Microsoft AI operates one of the world’s most advanced AI training infrastructures, featuring multi-gigawatt clusters spanning tens of thousands of high-performance GPUs, ultra-low-latency NVLink/NVSwitch networks, and innovative liquid-cooling systems. Our team is seeking a Member of Technical Staff, Hardware Health, to ensure these systems deliver sustained reliability, performance, and availability across exascale-class deployments. 

We work closely with research, hardware, datacenter, and platform engineering teams to develop predictive health models, failure detection frameworks, and autonomous remediation systems that keep our AI clusters operating at frontier scale. 

Our newly formed organization, Microsoft AI, is dedicated to advancing Copilot and other consumer AI products and research. The team is responsible for Copilot, Bing, Edge, and generative AI research. Join us and help shape the future of personal computing. 

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees, we embrace a growth mindset, innovate to empower others, and collaborate to achieve shared goals. Every day, we build on our values of respect, integrity, and accountability to foster a culture of inclusion where everyone can thrive at work and beyond. 

MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location. This expectation is subject to local law and may vary by jurisdiction. 



Responsibilities
  • Design and develop next-generation hardware health monitoring and diagnostic frameworks for large GPU clusters (NVL16/NVL72/GB200+ scale).
  • Build predictive analytics pipelines leveraging telemetry, power, and thermal data to anticipate hardware degradation and systemic issues.
  • Collaborate with silicon, firmware, and datacenter engineers to identify root causes and remediate large-scale hardware anomalies.
  • Define system health KPIs (e.g., NIS/RIS, MTBF, failure domain analysis) and integrate them into real-time observability platforms.
  • Lead incident triage for high-impact GPU, network, and cooling issues across distributed clusters.
  • Drive automation in health management to reduce manual intervention to the top 5% of anomalies.
  • Partner with cross-functional teams to influence hardware design for reliability, thermal efficiency, and serviceability.


Qualifications

Required Qualifications: 

  • Bachelor’s Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience.

Preferred Qualifications: 

  • Master’s Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor’s Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python 
    • OR equivalent experience.
  • Experience working with large-scale HPC or GPU systems (NVIDIA H100/GB200 or equivalent).
  • Deep understanding of GPU architecture, high-speed interconnects (NVLink, InfiniBand, RoCE), and large datacenter topologies.
  • Proficiency in hardware telemetry, diagnostics, or failure analysis tools.
  • Experience with exascale-class systems or cloud-scale AI clusters.
  • Familiarity with reliability modeling, machine learning-based anomaly detection, or predictive maintenance.
  • Contributions to large-scale infrastructure operations, supercomputing centers, or AI hardware design. 


Software Engineering IC5 – The typical base pay range for this role across the U.S. is USD $142,800 – $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 – $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay

Software Engineering IC6 – The typical base pay range for this role across the U.S. is USD $165,600 – $296,400 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $220,800 – $331,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.




Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar jobs

Senior Applied Scientist

Redmond, US
Applied Sciences

Principal Applied Scientist

Multiple Locations, China
Applied Sciences

Principal Applied Scientist

Suzhou, China
Applied Sciences

Multiple Locations, United States · New York, United States

Health Strategy & Commercialization

Location
Multiple Locations, United States · New York, United States
Job Number
200037172-en-2
City
Multiple Locations, New York
Team
Health
Country
United States
Discipline
Business Development
Overview

Microsoft AI Health

Help Build the Future of AI in Healthcare

Microsoft AI is building frontier AI products and platforms designed to transform how people experience healthcare. We are creating new capabilities that can improve access, support clinicians, empower patients, and unlock new models of care across the healthcare ecosystem. 

 


We are looking for an exceptional builder to help bring these innovations to market.
This role sits at the center of product strategy, healthcare ecosystem development, commercialization, partnerships, reimbursement, policy, and go-to-market execution. You will help shape how Microsoft AI Health turns breakthrough technology into real-world healthcare impact.
You will not be joining a mature business with a fixed playbook.


You will help create the playbook.


This is a rare opportunity for someone who thrives in ambiguity, thinks from first principles, and wants to help define what becomes possible when frontier AI meets healthcare.

 

The Opportunity

This is not a traditional partnerships role.
It is a full-stack builder role for someone who can move fluidly across strategy, product, partnerships, commercialization, and execution.
You will help answer some of the most important questions in AI and healthcare:

  • What new healthcare experiences become possible with frontier AI?
  • How should AI-powered healthcare products reach patients, providers, payers, and health systems?
  • Which partnerships can accelerate adoption and deepen product impact?
  • What business models and reimbursement pathways will support scalable growth?
  • How do we turn early product-market signals into durable, repeatable go-to-market motions?

The right person will bring the structured thinking of a strategist, the ownership mindset of an operator, the curiosity of a product builder, and the adaptability of someone who has built in startup-like environments.


What You’ll Do

Build and Commercialize New Healthcare Products

Help shape early-stage AI healthcare products from concept through market adoption. You will work across product, engineering, research, business development, policy, and commercial teams to define what should be built, how it should be positioned, and how it should scale.

Shape Go-To-Market Strategy

Develop commercialization strategies for new healthcare offerings, including product-market fit, customer segmentation, distribution, pricing, growth strategy, and sales enablement. You will help determine how Microsoft AI Health brings products to market in a way that is thoughtful, credible, and scalable.

Build Strategic Partnerships

Identify, develop, and execute partnerships across the healthcare ecosystem, including providers, payers, telehealth organizations, digital health companies, health systems, and other key industry stakeholders. You will help determine which partnerships are essential to product value, market adoption, and long-term impact.

Influence Product Direction

Bring market, customer, partner, and ecosystem insight directly into product strategy. You will help translate healthcare needs into product opportunities and ensure we are building solutions that matter to real users.

Shape Ecosystem and Reimbursement Strategy

Partner with government affairs, policy, legal, and reimbursement stakeholders to understand and influence the conditions required for scalable adoption. This may include exploring Medicare, Medicaid, commercial payer, employer, and enterprise pathways.

Create Leverage Across Microsoft

Work with existing Microsoft commercial teams, including MCAPS and other go-to-market partners, to determine when to build dedicated motions and when to enable broader Microsoft channels.



Responsibilities

Who You Are

You are a builder at heart.
You are energized by ambiguous problems, emerging markets, and the opportunity to create something new. You are comfortable moving from a high-level strategy conversation to a detailed product or partnership decision. You can work with executives, product leaders, engineers, policy experts, and external partners, while maintaining clarity, urgency, and ownership.
You do not need every answer upfront. You know how to form hypotheses, test them quickly, learn from evidence, and iterate.
You are low ego, high ownership, deeply curious, and mission-driven. You care about building products and businesses that can meaningfully improve people’s lives.


What Great Looks Like

The strongest candidates will demonstrate:

Entrepreneurial Ownership

You take responsibility end to end. You do not wait for perfect structure, perfect information, or permission to move important work forward.

Strategic Rigor

You bring clear thinking to ambiguous problems. You can separate signal from noise, develop hypotheses, evaluate tradeoffs, and make strong recommendations.

Healthcare Fluency

You understand the complexity of the U.S. healthcare ecosystem, including providers, payers, reimbursement, health systems, digital health, policy, and commercial adoption.

Commercial Instinct

You understand how products become businesses. You can identify market opportunities, define growth motions, evaluate business models, and create paths to scale.

Product Orientation

You think deeply about users, customer needs, product value, and adoption. You can influence product decisions without needing to formally own the product roadmap.

Partnership Judgment

You know how to identify the right partners, structure meaningful collaborations, and build relationships that create durable value.

Adaptability

You are comfortable operating like a “stem cell” talent, able to flex across disciplines and specialize as the business evolves.

Clear Communication

You simplify complexity. You communicate with precision. You create alignment across senior stakeholders and cross-functional teams.


MAI Culture Alignment

Microsoft AI is building with a high bar for talent, speed, rigor, and impact. The people who thrive here are deeply mission-driven, intellectually honest, low ego, and energized by doing the best work of their lives.
This role requires someone who embodies that culture.
We are looking for someone who:

  • Raises the talent density of every team they join
  • Operates with urgency without sacrificing long-term judgment
  • Takes ownership through completion
  • Uses evidence, customer insight, and clear reasoning to guide decisions
  • Asks hard questions and updates their thinking quickly
  • Simplifies complexity rather than adding process
  • Works as a genuine team player with low ego
  • Communicates clearly, directly, and thoughtfully
  • Builds for users, not internal convenience
  • Thrives in small, high-impact teams with minimal hierarchy

This is a role for someone who wants to help build, learn, iterate, and create impact at the frontier of AI and healthcare.


Ideal Background

We are less focused on titles and more focused on pattern recognition.
Strong candidates may bring a combination of:

  • Top-tier strategy consulting experience in healthcare
  • Operating experience in healthcare technology, digital health, AI, life sciences, payer, provider, or healthcare services
  • Startup or high-growth company experience
  • Product commercialization experience
  • Strategic partnerships or business development experience
  • New venture creation, incubation, or zero-to-one building experience
  • Experience working across policy, reimbursement, and healthcare commercialization

Relevant backgrounds may include candidates from firms such as McKinsey, Bain, BCG, LEK, Oliver Wyman, Accenture Strategy, or other high-caliber healthcare strategy environments, particularly those who have moved into operating roles and helped build products, businesses, partnerships, or growth engines.


Why This Role Matters

Most healthcare roles optimize existing systems.
Most AI roles focus on building technology.
This role is about connecting the two.
You will help translate frontier AI into healthcare products, partnerships, business models, and go-to-market strategies that can create real-world impact. You will work on problems where the answers are not obvious yet, but where the opportunity is enormous.
For the right person, this is more than a partnerships role.
It is a chance to help define the future of AI in healthcare, from the ground up.



Qualifications

Required/minimum qualifications

  • Bachelor’s Degree in Business, Liberal Arts, Sciences, or related field AND 12+ years relevant work experience in a relevant role in healthcare (e.g. healthcare startup, consulting, venture capital, business development, or related field) OR equivalent experience.

Additional or preferred qualifications

  • Deep knowledge of the US healthcare market
  • Exceptional ability to develop and build trusting relationships with executive-level partners
  • Strong communication, presentation and project management skills
  • Ability to produce high quality written and presentation materials (PowerPoint, Word) with excellent attention to detail
  • Strong numerical analysis (Excel) and financial literacy (understand fundamentals of P&L, balance sheet etc)
  • Drive and initiative to maintain momentum, even when working solo
  • Work at high speed with autonomy in a rapidly changing environment
  • Relish ambiguity/uncertainty


Business Development IC5 – The typical base pay range for this role across the U.S. is USD $130,900 – $251,900 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $165,600 – $272,300 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay

Business Development IC6 – The typical base pay range for this role across the U.S. is USD $155,800 – $277,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $202,400 – $303,600 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.




Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar jobs

Senior Applied Scientist

Redmond, US
Applied Sciences

Principal Applied Scientist

Multiple Locations, China
Applied Sciences

Principal Applied Scientist

Suzhou, China
Applied Sciences

Multiple Locations, United States

Principal Product Manager

Location
Multiple Locations, United States
Job Number
200048525-en-1
City
Multiple Locations
Team
MAI
Country
United States
Discipline
Product Management
Overview

This role is part of the Microsoft AI Team. The MAI Team mission is to empower every person and organization on the planet to achieve more. As team members, we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals.  

MAI is a startup-like team inside Microsoft, created to push the boundaries of AI toward Humanist Superintelligence — ultra-capable systems that remain controllable, safety-aligned, and anchored to human values. Our mission is to create AI that amplifies human potential while ensuring humanity remains firmly in control. We aim to deliver breakthroughs that benefit society — advancing science, education, and global well-being.

Our team is led by Mustafa Suleyman, CEO of Microsoft AI. We bring together leading researchers, engineers, and product builders who have contributed to many of the most significant advances in modern AI.

As a Principal Product Manager, you will help shape the next generation of AI models. Working backward from customer and product needs, you will identify opportunities for new research, develop evaluations and datasets, and partner closely with researchers and engineers to drive model development and deployment. You will play a critical role in translating emerging capabilities into impactful products and experiences.

 

Candidates that thrive at MAI 

  • Hold an exceptionally high bar for delivering excellence in every task  

  • Are able to balance speed and quality even when delivering ambiguous or emerging goals  

  • Bring ambitious bias for action, push boundaries, and take ownership   

  • Keep things simple  

  • Deeply and authentically value teamwork  

  • Maintain an optimistic, humble, can-do attitude to all tasks, big and small 



Responsibilities
  • Own and deliver the model roadmap: Define the roadmap in collaboration with key stakeholders. Understand and navigate tradeoffs quickly and strategically. Balance quality with speed. Drive continuous progress through ambitious and energized product leadership.   

  • Partner to build alignment: Work closely with partner teams to understand and translate their product needs into model requirements.  

  • Define model evaluations: Take a scientific, data-driven approach to product success. Build relevant benchmarks that enable us to climb to frontier quality on the tasks that matter for the product.   

  • Determine launch readiness: Establish clear criteria for model launches. Own the go/no-go decisions that balance capability, risk, and product goals. Ensure strong alignment and transparent, regular communication throughout the launch process. 

  • Stay at the frontier: Track the rapidly evolving AI landscape. Translate what you learn into actionable product priorities. 



Qualifications

Required/Minimum Qualifications

  • Bachelor’s Degree AND 8+ years experience in product/service/program management or software development
    • OR equivalent experience.

Additional/Preferred Qualifications

  • Bachelor’s Degree AND 12+ years experience in product/service/program management or software development
    • OR equivalent experience.
  • 4+ years experience taking a product, feature, or experience to market (e.g., design, addressing product market fit, and launch, internal tool/framework).
  • 6+ years experience improving product metrics for a product, feature, or experience in a market (e.g., growing customer base, expanding customer usage, avoiding customer churn).
  • 6+ years experience disrupting a market for a product, feature, or experience (e.g., competitive disruption, taking the place of an established competing product).
  • 5+ experience defining product requirements for AI models and driving decisions and execution across multiple senior level stakeholders.
  • 5+ years direct, hands-on collaboration with model researchers and ML engineers in areas such as model training, evaluation, computing infrastructure, agentic and reasoning models, where you can point to your concrete contribution.  
  • Proven ability to ship best-in-class AI models or AI enabled products in a fast-paced, competitive environment.
  • Demonstrated ability to understand and reason with precision through the technical and commercial drivers required to determine a Go to Market (GTM) strategy for new AI products. 
  • Proven ability to bring successful models, agents, or native AI products to market by deeply engaging with customers and developers across different verticals. 


Product Management IC5 – The typical base pay range for this role across the U.S. is USD $142,800 – $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 – $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay

Product Management IC6 – The typical base pay range for this role across the U.S. is USD $165,600 – $296,400 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $220,800 – $331,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.




Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar jobs

Senior Applied Scientist

Redmond, US
Applied Sciences

Principal Applied Scientist

Multiple Locations, China
Applied Sciences

Principal Applied Scientist

Suzhou, China
Applied Sciences

Multiple Locations, United States

Member of Technical Staff – AI Data Platform, Frontier Models

Location
Multiple Locations, United States
Job Number
200051482-en-1
City
Multiple Locations
Team
MAI AI Platform
Country
United States
Discipline
Data Engineering
Overview
Help build the world’s most advanced multimodal dataset at Microsoft AI
We are on a mission to create the largest and most advanced multimodal dataset in the world. This dataset, spanning all modalities from across the web and beyond, will power the training of the world’s most capable AI frontier models, pushing the boundaries of scale, performance, and product deployment.

The AI Data Infra team at Microsoft AI is responsible for building data infrastructure to help MAI teams to generate the biggest and best training dataset. Our work involves data pipelines, Spark, Ray, Vector Databases, and all other aspects of data infra.
We are looking for outstanding individuals excited about contributing to the next generation of systems that will transform the field. In particular, we are looking for candidates who:

Are passionate about the role of data in large-scale AI model training
Will thrive in a highly collaborative, fast-paced environment
Have a high degree of expertise and pay close attention to details
Demonstrate a proactive attitude and enthusiasm for exploring new methods and technologies
Effectively manage multiple responsibilities and can adjust to shifting priorities.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
 
Starting January 26, 2026, MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location. This expectation is subject to local law and may vary by jurisdiction.
 
This role is part of Microsoft AI’s Superintelligence Team. The MAIST is a startup-like team inside Microsoft AI, created to push the boundaries of AI toward Humanist Superintelligence—ultra-capable systems that remain controllable, safety-aligned, and anchored to human values. Our mission is to create AI that amplifies human potential while ensuring humanity remains firmly in control. We aim to deliver breakthroughs that benefit society—advancing science, education, and global well-being.
 
We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models! 


Responsibilities

Build AI Data Infrastructure and Data Engines
Develop and evolve large-scale AI data infrastructure for frontier AI lab. Own data collection, ingestion, cleaning, curation, generation, governance, metadata management, query , analytic and hybrid search , building reusable, observable, and explainable EB-scale data platform that support high-quality data for pre-training and post-training workloads.

Develop Intelligent Multimodal Data Processing Systems
Lead automated understanding and processing of text, images, video, documents. Build labeling and taxonomy systems, semantic feature extraction, data quality modeling, and automated data governance capabilities. Design and train models for classification, recognition, captioning, quality scoring, prediction, and related data-processing tasks.

Build AI-Native Data Pipelines
Leverage LLMs, VLMs, and Agents to build intelligent data pipelines that automate data collection, filtering, deduplication, quality diagnosis, annotation/re-labeling, generation, scheduling, orchestration, and anomaly detection. Use AI-native workflows to significantly reduce manual data operations and improve pipeline scalability and efficiency.

Build AI-Native Data Storage and Table Layers
Design open, AI-optimized storage using Lance, Iceberg, Paimon, and Parquet. Support multimodal
data and embeddings, fast random access and scans, schema evolution, transactions, versioning/
time travel, indexing, and interoperable access across training and query engines.

Discover and Build Rare, High-Value Datasets
Develop differentiated datasets for challenging domains, including web/PDF/encyclopedic/private/query-based text data as well as real-world multimodal scenarios such as retail inspection, warehouse inspection, healthcare, education, OCR, GUI interaction, and multimodal trajectories. Apply a combination of real-world data collection, human annotation, and synthetic data generation to produce rare and high-value datasets.

Drive the Data–Model–Evaluation Iteration Loop
Use evaluation feedback and model failure analysis to identify data gaps and design targeted datasets and synthetic-data strategies. Establish observable metrics that quantify the contribution of data to model capability improvements, continuously update datasets during training, and build an iterative Data → Model → Evaluation → Data optimization loop.

Partner Closely with Model and Training Teams
Collaborate with model researchers and training engineers on training-data construction, data feedback loops, targeted data mining, and evaluation-driven iteration. Use data as a primary lever for improving model capability and become a core driver of AI model training performance.



Qualifications
  • Required:
    Master’s Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 3+ years experience in business analytics, data science, software development, data modeling, or data engineering OR Bachelor’s Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering OR equivalent experience. 
    Experience with distributed data Platforms such as Spark, Flink or Ray.
    Proficient experience in Python and experience with SQL and Shell. 
    Experience with Multimodal Data (Text, Image, Video, or Audio)

    Preferred:
  • Hands-on experience building datasets for LLM/VLM/multimodal model pre-training or post-training.
  • Experience with synthetic data, including text generation, image-text synthesis, rendering/diffusion-based generation, or multimodal trajectory generation for GUI, search, file-operation, or agentic tasks.
  • Familiarity with major evaluation benchmarks such as MMLU, MMBench, MM-BrowseComp, with practical experience using evaluation results and failure analysis to drive targeted data improvements.
  • Experience with open file and table formats such as Lance, Iceberg, Paimon and Parquet,including schema evolution, versioning, transactions, indexing, and performance optimization for multimodal AI workloads.
  • Solid understanding of LLMs, speech/audio models, vision models, and multimodal models. Hands-on experience with large-scale AI data construction, cleaning, synthesis, or quality evaluation.
  • Experience building ETL systems, data models, data pipelines, or data warehouses is strongly preferred. Experience processing large-scale text, image, or video datasets is a plus.
  • Familiarity with Agents and modern LLM toolchains, with practical experience—or strong interest—in applying LLMs to data production, analysis, quality control, governance, and pipeline automation.


Data Engineering IC5 – The typical base pay range for this role across the U.S. is USD $142,800 – $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 – $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay

Data Engineering IC6 – The typical base pay range for this role across the U.S. is USD $165,600 – $296,400 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $220,800 – $331,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.




Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar jobs

Senior Applied Scientist

Redmond, US
Applied Sciences

Principal Applied Scientist

Multiple Locations, China
Applied Sciences

Principal Applied Scientist

Suzhou, China
Applied Sciences

Multiple Locations, United States

Member of Technical Staff – Data Flywheel Infra, Frontier Models

Location
Multiple Locations, United States
Job Number
200051479-en-1
City
Multiple Locations
Team
MAI AI Platform
Country
United States
Discipline
Data Engineering
Overview

We are looking for a Data Flywheel Infrastructure Engineer to build the infrastructure that continuously turns 1P data, 3P data, model signals, evaluation results, and synthetic data into high-quality training data for frontier LLM and multimodal models.
This role owns the systems connecting:
Data Acquisition → Governance & Compliance → Curation → Training → Evaluation → Failure Mining → Data Improvement
A critical part of the role is enabling aggressive data iteration while ensuring that every dataset is secure, policy-compliant, rights-aware, traceable, and auditable.

Starting January 26, 2026, MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location. This expectation is subject to local law and may vary by jurisdiction.

This role is part of Microsoft AI’s Superintelligence Team. The MAIST is a startup-like team inside Microsoft AI, created to push the boundaries of AI toward Humanist Superintelligence—ultra-capable systems that remain controllable, safety-aligned, and anchored to human values. Our mission is to create AI that amplifies human potential while ensuring humanity remains firmly in control. We aim to deliver breakthroughs that benefit society—advancing science, education, and global well-being.
 
We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models! 


Responsibilities
  1. Build 1P & 3P Data Flywheel Infrastructure
    Build scalable systems for ingesting, processing, curating, versioning, and serving first-party and third-party data for pre-training and post-training. Connect model failures, evaluations, and product signals back into targeted data acquisition, generation, and improvement workflows.
  2. Own Data Governance, Security & Compliance Infrastructure
    Build governance and policy enforcement directly into the data platform, including:
    • Data provenance and lineage
    • Usage rights, licensing, and consent metadata
    • PII / sensitive-data detection and protection
    • Access control and data isolation
    • Retention and deletion enforcement
    • Geographic and regulatory restrictions
    • Dataset approval and audit workflows
    • Training eligibility and purpose-based usage controls
  1. Build Policy-Aware Data Acquisition & Curation Systems
    Develop automated pipelines for 1P and 3P data ingestion, classification, filtering, deduplication, quality scoring, semantic enrichment, and dataset construction.
    Make governance policies machine-enforceable so that data can automatically be included, excluded, quarantined, or restricted based on its origin, license, sensitivity, consent, geography, and intended model use.
  2. Build Evaluation-to-Data Feedback Loops
    Convert model evaluations and real-world failure signals into actionable data tasks through failure clustering, hard-example mining, long-tail discovery, capability-gap detection, and targeted dataset generation.
    Enable rapid iteration from:
    Model Failure → Data Gap → Data Intervention → Training → Evaluation
  3. Build Synthetic & AI-Native Data Pipelines
    Use LLMs, VLMs, and Agents to automate data generation, labeling, filtering, quality validation, enrichment, and transformation.
    Maintain clear provenance between human-created, first-party, third-party, model-generated, and derived data, and enforce appropriate policies across each category.
  4. Build Data Quality, Attribution & Observability
    Develop metrics and infrastructure to measure dataset quality, coverage, diversity, contamination, duplication, policy compliance, and contribution to model capability improvements.
    Enable researchers to understand which data improves which capabilities and under what governance constraints.


Qualifications

Required

•  Master’s Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 3+ years experience in business analytics, data science, software development, data modeling, or data engineering OR Bachelor’s Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering OR equivalent experience.  

•  Software Engineering experience using Python,SQL, Spark/Flink/Ray

Preferred

  • Experience building AI training-data governance platforms, including provenance, licensing/rights metadata, consent management, PII handling, policy enforcement, or auditable lineage.
  • Experience managing third-party datasets, data partnerships, licensed content, or externally sourced data with complex contractual and usage restrictions.
  • Experience building privacy- and security-aware systems for first-party product or user data, including isolation, access controls, retention/deletion, and purpose limitation.
  • Experience with data clean rooms, privacy-preserving processing, de-identification, confidential computing, or secure data collaboration.
  • Experience building evaluation → failure mining → data generation → training feedback loops.
  • Experience with synthetic data, model graders, reward signals, hard-example mining, active learning, or data-mixture optimization.
  • Experience with multimodal or agentic datasets including text, image, video, audio, web, GUI, tool-use, or interaction trajectories.
  • Understanding of Modern LLM training workflows including Pre-training, SFT, RL/post-training, evaluation, and synthetic data. 
  • Strong understanding of data governance, security, privacy, provenance, access control, and data lifecycle management.





Data Engineering IC4 – The typical base pay range for this role across the U.S. is USD $119,800 – $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 – $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay

Data Engineering IC5 – The typical base pay range for this role across the U.S. is USD $142,800 – $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 – $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.




Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar jobs

Senior Applied Scientist

Redmond, US
Applied Sciences

Principal Applied Scientist

Multiple Locations, China
Applied Sciences

Principal Applied Scientist

Suzhou, China
Applied Sciences

Multiple Locations, United States

Member of Technical Staff – Data Flywheel Infra, Frontier Models

Location
Multiple Locations, United States
Job Number
200052218-en-1
City
Multiple Locations
Team
MAI AI Platform
Country
United States
Discipline
Data Engineering
Overview
We are looking for a Data Flywheel Infrastructure Engineer to build the infrastructure that continuously turns 1P data, 3P data, model signals, evaluation results, and synthetic data into high-quality training data for frontier LLM and multimodal models.
This role owns the systems connecting:
Data Acquisition → Governance & Compliance → Curation → Training → Evaluation → Failure Mining → Data Improvement
A critical part of the role is enabling aggressive data iteration while ensuring that every dataset is secure, policy-compliant, rights-aware, traceable, and auditable.
 
Starting January 26, 2026, MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location. This expectation is subject to local law and may vary by jurisdiction.
 
This role is part of Microsoft AI’s Superintelligence Team. The MAIST is a startup-like team inside Microsoft AI, created to push the boundaries of AI toward Humanist Superintelligence—ultra-capable systems that remain controllable, safety-aligned, and anchored to human values. Our mission is to create AI that amplifies human potential while ensuring humanity remains firmly in control. We aim to deliver breakthroughs that benefit society—advancing science, education, and global well-being.
 
We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models! 


Responsibilities

Build 1P & 3P Data Flywheel Infrastructure
Build scalable systems for ingesting, processing, curating, versioning, and serving first-party and third-party data for pre-training and post-training. Connect model failures, evaluations, and product signals back into targeted data acquisition, generation, and improvement workflows.

Own Data Governance, Security & Compliance Infrastructure
Build governance and policy enforcement directly into the data platform, including:

    • Data provenance and lineage
    • Usage rights, licensing, and consent metadata
    • PII / sensitive-data detection and protection
    • Access control and data isolation
    • Retention and deletion enforcement
    • Geographic and regulatory restrictions
    • Dataset approval and audit workflows
    • Training eligibility and purpose-based usage controls

Build Policy-Aware Data Acquisition & Curation Systems
Develop automated pipelines for 1P and 3P data ingestion, classification, filtering, deduplication, quality scoring, semantic enrichment, and dataset construction.
Make governance policies machine-enforceable so that data can automatically be included, excluded, quarantined, or restricted based on its origin, license, sensitivity, consent, geography, and intended model use.

Build Evaluation-to-Data Feedback Loops
Convert model evaluations and real-world failure signals into actionable data tasks through failure clustering, hard-example mining, long-tail discovery, capability-gap detection, and targeted dataset generation.
Enable rapid iteration from:
Model Failure → Data Gap → Data Intervention → Training → Evaluation

Build Synthetic & AI-Native Data Pipelines
Use LLMs, VLMs, and Agents to automate data generation, labeling, filtering, quality validation, enrichment, and transformation.
Maintain clear provenance between human-created, first-party, third-party, model-generated, and derived data, and enforce appropriate policies across each category.

Build Data Quality, Attribution & Observability
Develop metrics and infrastructure to measure dataset quality, coverage, diversity, contamination, duplication, policy compliance, and contribution to model capability improvements.
Enable researchers to understand which data improves which capabilities and under what governance constraints.

 



Qualifications

Required

  • Master’s Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 4+ years experience in business analytics, data science, software development, data modeling, or data engineering OR Bachelor’s Degree in Computer Science, Math, Software Engineering, Computer Engineering, or related field AND 6+ years experience in business analytics, data science, software development, data modeling, or data engineering OR equivalent experience.  
  • Software Engineering experience using Python,SQL, Spark/Flink/Ray

    Preferred

  • Experience building AI training-data governance platforms, including provenance, licensing/rights metadata, consent management, PII handling, policy enforcement, or auditable lineage.
  • Experience managing third-party datasets, data partnerships, licensed content, or externally sourced data with complex contractual and usage restrictions.
  • Experience building privacy- and security-aware systems for first-party product or user data, including isolation, access controls, retention/deletion, and purpose limitation.
  • Experience with data clean rooms, privacy-preserving processing, de-identification, confidential computing, or secure data collaboration.
  • Experience building evaluation → failure mining → data generation → training feedback loops.
  • Experience with synthetic data, model graders, reward signals, hard-example mining, active learning, or data-mixture optimization.
  • Experience with multimodal or agentic datasets including text, image, video, audio, web, GUI, tool-use, or interaction trajectories.
  • Understanding of Modern LLM training workflows including Pre-training, SFT, RL/post-training, evaluation, and synthetic data. 
  • Strong understanding of data governance, security, privacy, provenance, access control, and data lifecycle management.


Data Engineering IC5 – The typical base pay range for this role across the U.S. is USD $142,800 – $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 – $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay

Data Engineering IC6 – The typical base pay range for this role across the U.S. is USD $165,600 – $296,400 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $220,800 – $331,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.




Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar jobs

Senior Applied Scientist

Redmond, US
Applied Sciences

Principal Applied Scientist

Multiple Locations, China
Applied Sciences

Principal Applied Scientist

Suzhou, China
Applied Sciences
English (United States)
Your Privacy Choices Opt-Out Icon Your Privacy Choices
Consumer Health Privacy Sitemap Contact Microsoft Privacy Manage cookies Terms of use Trademarks Safety & eco Recycling About our ads