Overview
We are looking for a principal engineer to build the monitoring, detection, automation, and response capabilities that protect Microsoft AI.
Microsoft AI develops and trains first-party models that Microsoft publishes and uses across products including Foundry, Copilot, Microsoft 365, and MDASH. The work spans model development and serving, large-scale compute, shared infrastructure, engineering systems, identity, data, and productivity services - the whole operational system that makes large-scale training, experimentation, and deployment possible. In the last year, MAI evolved from a research lab to a full production model factory. This role is designed for an engineer who has built or substantially transformed security monitoring and response capabilities and is prepared to do so again. You will set technical direction and build production systems across telemetry collection, security data, SIEM, detection engineering, hunting, SOAR, incident response, and AI-native security operations.
This role exists to build a function, not simply administer an existing SOC stack. You will help define what Microsoft AI must observe, which telemetry is collected and in what order, how detections are created and evaluated, how investigations are assembled, and where automated systems can safely recommend or take action. You will deploy and manage collection agents such as AzSecPack or comparable modern sensors, shape the SIEM and security-data architecture, and replace repeated manual work with reliable automation.
You will also bring patterns emerging in AI safety into security operations: continuously running evaluators or judges, evidence-grounded analysis, automated alerting, and agents that can investigate or respond within explicit authority. These systems must be engineered rather than merely prompted. They need evaluation sets, quality measures, scoped identities and permissions, traceable evidence, deterministic guardrails, human escalation, safe rollback, and clear rules for when automation may act. This entire category is new – we don’t expect you to be an expert on day one – but it’s helpful if you are (or can do some research before we chat).
This is a principal individual-contributor role that leads through architecture, code, technical judgment, prototypes, production ownership, and influence across a growing security-engineering team and partner organizations. MAI Security earns trust by staying close to the systems it protects and remaining accountable for detection and response outcomes, not just for raising risks. We build working systems, make coverage and residual risk explicit, and are honest about what our detections and agents can and cannot yet be trusted to do. The goal is to help Microsoft AI move quickly on frontier work by making security risk visible and manageable, creating safer options rather than new bottlenecks.
What You Will Do
Build the Security Telemetry Foundation
- Define a security-observability strategy across identities, endpoints, hosts, workloads, containers, orchestration layers, networks, cloud control planes, developer systems, model infrastructure, data systems, and business applications.
- Define telemetry and detection requirements for model and agent access, tool invocation, training and evaluation jobs, model and data artifacts, safety signals relevant to cyber defense, and anomalous use of high-consequence capabilities.
- Deploy, integrate, and operate collection agents such as AzSecPack or modern equivalents. Establish rollout, upgrade, health, tamper-resistance, coverage, performance, and incident-use requirements.
- Design reliable telemetry pipelines with clear schemas, enrichment, routing, retention, access control, data-quality checks, lineage, cost controls, and resilience to partial failure or adversarial interference.
- Identify high-consequence visibility gaps, work with system owners to close them, and make monitoring requirements part of platform and service architecture.
- Establish measures for coverage, freshness, completeness, health, query performance, detection dependencies, and operational value.
Build the SIEM, Detection, and Response Platform
- Build detections and hunting capabilities grounded in threat models, attack paths, threat intelligence, adversary simulation, incidents, and frontier-model infrastructure.
- Engineer detection lifecycle management: version control, testing, deployment, coverage mapping, quality measurement, tuning, ownership, retirement, and validation against representative data.
- Build or integrate SOAR capabilities for enrichment, triage, evidence gathering, case management, containment, recovery, and corrective action.
- Ensure monitoring and response capabilities remain operable during incidents involving compromised identities, token misuse, telemetry tampering, data exfiltration, provider compromise, or compromise of training and model infrastructure.
Create AI-Native Security Operations
- Build continuously running evaluators or judges that assess security events, system behavior, control health, and emerging risk using grounded evidence and explicit decision criteria.
- Design agents that can generate and run queries, correlate evidence, form and test hypotheses, summarize investigations, recommend actions, and execute narrowly scoped response steps when confidence and authority permit.
- Create evaluation sets, adversarial tests, replay environments, quality measures, and human-review mechanisms. Measure false positives, false negatives, calibration, latency, cost, and operational usefulness.
- Combine deterministic logic with model-based reasoning deliberately. Avoid opaque automation where evidence, repeatability, or safety is insufficient.
- Apply lessons from Microsoft AI safety-monitoring systems while accounting for the different adversaries, failure modes, authority, and evidentiary requirements of cyber defense.
Lead Technical Strategy, Decisions, and Adoption
- Build working prototypes and production systems in code. Debug pipelines, agents, detections, integrations, and incidents directly.
- Create a multi-year technical roadmap while delivering useful increments early. Make dependencies, tradeoffs, ownership, and residual risk explicit.
- Partner with research, platform, infrastructure, developer-experience, product, safety, privacy, and central security teams to integrate telemetry and response into the systems they own.
- Help accountable owners make risk-informed decisions on new architectures, exceptions, and access patterns. Provide practical mitigations and reusable precedent rather than a pass/fail gate.
- Establish reusable architectures, libraries, interfaces, and standards that allow other teams to contribute detections, data sources, automations, and agent capabilities safely.
- Escalate blocked telemetry, coverage, or automation-authority decisions with evidence, options, and a clear decision request.
- Mentor engineers and incident responders, lead difficult technical reviews, and represent the function in high-consequence architecture, incident, and investment decisions.
What Makes This Role Different
- You are building the function. You will define the architecture, operating model, engineering standards, and roadmap rather than tune an inherited queue.
- AI is part of the operating system, not a chat interface. Judges and agents must be evaluated, permissioned, observable, auditable, and safe enough for production security work.
- The scope spans classic and emerging security operations. Host agents, SIEM, SOAR, hunting, detections, and incident response must work alongside monitoring for model, training, data, and agentic systems.
- You will build and operate. Principal-level influence comes from systems that work in production.
- You can create Microsoft-scale leverage. The role can use and influence extensive security platforms, threat intelligence, identity, cloud, and research capabilities while retaining accountability for Microsoft AI outcomes.
What Success Looks Like
- Microsoft AI has a coherent security-observability architecture with known coverage, accountable owners, measurable health, and prioritized gaps.
- High-value systems have reliable collection agents and telemetry pipelines that remain trustworthy and useful during incidents.
- The SIEM and detection platform supports repeatable engineering, testing, deployment, quality measurement, hunting, and investigation.
- Repeated investigation and response work has been automated safely, with clear evidence, authority, escalation, and rollback.
- AI judges and agents demonstrate measurable operational value against representative and adversarial evaluation sets.
- Researchers, engineers, and responders can understand why an alert or automated action occurred and can challenge, improve, or disable it.
- Incidents and exercises lead to stronger telemetry, detections, automations, and platform controls.
Required/Minimum Qualifications
- Doctorate in Statistics, Mathematics, Computer Science, or related field AND 3+ years experience in software development lifecycle, large-scale computing, threat modeling, cyber security, anomaly detection, Security Operations Center (SOC) detection, threat analytics, security incident and event management (SIEM), information technology (IT), or operations incident response
- OR Master's Degree in Statistics, Mathematics, Computer Science, or related field AND 4+ years experience in software development lifecycle, large-scale computing, threat modeling, cyber security, anomaly detection, Security Operations Center (SOC) detection, threat analytics, security incident and event management (SIEM), information technology (IT), or operations incident response
- OR Bachelor's Degree in Statistics, Mathematics, Computer Science, or related field AND 6+ years experience in software development lifecycle, large-scale computing, threat modeling, cyber security, anomaly detection, Security Operations Center (SOC) detection, threat analytics, security incident and event management (SIEM), information technology (IT), or operations incident response
- OR equivalent experience.
Preferred/Additional Qualifications
Preferred qualifications are not independent pass/fail gates; candidates who meet some but not all are encouraged to apply.
- Doctorate in Statistics, Mathematics, Computer Science, or related field AND 5+ years experience in software development lifecycle, large scale computing, threat modeling, cyber security, or anomaly detection
- OR Master's Degree in Statistics, Mathematics, Computer Science, or related field AND 8+ years experience in software development lifecycle, large scale computing, threat modeling, cyber security, or anomaly detection
- OR Bachelor's Degree in Statistics, Mathematics, Computer Science, or related field AND 12+ years experience in software development lifecycle, large scale computing, threat modeling, cyber security, or anomaly detection
- OR equivalent experience.
- CISSP CISA CISM SANS OSCP Security+
- Extensive experience designing, building, and operating security-monitoring, detection, or response systems in complex production environments.
- Deep prior ownership in at least one of telemetry and collection, SIEM and detection engineering, or SOAR and response automation, with working breadth across several related domains.
- Strong software-engineering experience designing, debugging, testing, deploying, and operating production code and distributed systems.
- Experience establishing technical direction across multiple teams and turning an ambiguous security problem into an adopted platform or capability.
- Experience designing telemetry and detection systems for reliability, data quality, access control, observability, cost, scale, and adversarial conditions.
- Experience leading significant incidents or building the technical capabilities used to investigate, contain, and recover from them.
- Direct experience building and evaluating at least one production or near-production system using statistical, machine-learning, large-language-model, or agentic techniques for an operational or security use case, or equivalent evidence of rigorous capability in this area.
- Ability to explain architecture, evidence, uncertainty, tradeoffs, and recommendations to engineers and senior leaders.
- Experience securing AI or machine-learning infrastructure, large-scale compute, model-development systems, sensitive research environments, or high-value intellectual property.
- Experience building AI-assisted detection, investigation, or response systems with evaluation sets, human oversight, scoped permissions, and production safety mechanisms.
- Experience with Azure security telemetry, Microsoft Sentinel, Defender platforms, Kusto Query Language, AzSecPack, or comparable technologies.
- Experience designing agent or sensor deployment across heterogeneous hosts, containers, clusters, cloud environments, or specialized compute.
- Experience with adversary simulation, purple teaming, threat-informed defense, or detection coverage mapping.
- Experience creating a new security-operations capability where ownership, architecture, and processes were not already settled.
- Evidence of technical influence beyond one team through platforms, standards, open-source work, publications, incident leadership, or widely adopted engineering practices.
Security Operations Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.