Noida, India · Bengaluru, India · Hyderabad, India

Senior Principal Architect : Ads Trust & Safety AI Platform

Location
Noida, India · Bengaluru, India · Hyderabad, India
Job Number
200039959-en-3
City
Noida, Bengaluru, Hyderabad
Team
Content Engineering
Country
India
Discipline
Software Engineering
Overview

Microsoft Advertising serves billions of ad decisions every day across search, native, shopping, display, audience, and AI-powered advertising experiences. Trust & Safety is foundational to this marketplace: protecting users from harmful, deceptive, and low-quality content; protecting advertisers, publishers, and partners from fraud and abuse; and ensuring Microsoft Ads operates with strong policy, privacy, security, compliance, and regulatory integrity.We are seeking a deeply technical Senior Principal Architects to define and drive the next generation of the Ads Trust & Safety AI Platform. This is a engineering leadership role that entails setting technical direction, shaping platform strategy, influencing multiple engineering and science teams, and building durable systems that compound across Trust & Safety, Risk, Fraud, Security, Policy, Measurement, and Enforcement. This role is for a hands-on technical leader who can combine large-scale platform engineering with modern AI systems: agentic investigation workflows, LLM-powered reasoning, retrieval and evidence generation, heterogeneous model serving, risk intelligence, human-in-the-loop review, and high-integrity decisioning.The platform will power how Microsoft Ads understands advertisers, domains, landing pages, websites, business entities, policy risk, fraud signals, and adversarial behavior. It will enable faster and more reliable enforcement decisions, richer evidence for reviewers and investigators, safer automation, and stronger marketplace protection. A key part of this role is collaboration across the broader Trust ecosystem: Microsoft Ads, Microsoft Trust and Safety teams across products, Security,  Identity, and relevant worldwide industry peers and partners. The role requires someone who can evangelize technical direction, learn from adjacent teams and industry practices, align standards, and convert shared learnings into production-grade platform capabilities.



Responsibilities

AI Platform Architecture and Technical Strategy

  • Define the long-term architecture for the Ads Trust & Safety AI Platform across ingestion, signal acquisition, entity intelligence, retrieval, model orchestration, agentic workflows, decisioning, enforcement, human review, audit, and measurement.
  • Translate broad Trust & Safety, Risk, Fraud, Security, and Policy needs into reusable platform capabilities.
  • Establish reference architectures, design principles, technical standards, and engineering patterns for high-integrity AI and decisioning systems.
  • Drive architecture choices across latency, throughput, quality, cost, explainability, governance, reliability, and operational safety.
  • Identify critical platform gaps and create a roadmap that balances near-term delivery with long-term leverage.

Deep Research Agents and AI-Assisted Investigation

  • Architect platform for Deep Research Agents that investigate domains, landing pages, advertisers, business entities, ownership patterns, web presence, reputation, policy risk, and fraud signals.
  • Architect workflows that combine retrieval, crawling, structured evidence extraction, LLM reasoning, policy grounding, risk scoring and human in the loop.
  • Architect guardrails for agentic systems, including source provenance, confidence scoring, hallucination controls, audit logs, escalation paths, and human override.
  • Partner with Applied Science to convert AI research prototypes into production systems with clear quality, latency, cost, reliability, safety, and governance targets.

Entity Intelligence, Decisioning, and Enforcement

  • Architect systems for high-fidelity understanding of domains, websites, landing pages, advertisers, business identities, ownership structures, relationship graphs, reputation, and provenance.
  • Design real-time, nearline, and batch scoring systems for policy enforcement, fraud detection, abuse prevention, advertiser risk scoring, and marketplace protection.
  • Evolve abstractions for model orchestration, feature lookup, signal stores, retrieval, model versioning, decision logging, policy controls, fallbacks, and experimentation.

Security, Risk, and Fraud Platform

  • Architect systems to detect , learn and mitigate adversarial behavior across the advertiser lifecycle, including account creation, login events, payment changes, budget changes, campaign edits, creative changes, landing-page changes, and enforcement history.
  • Build sequential and event-based risk systems that reason over advertiser behavior over time rather than treating each decision as an isolated event.

Cross-Microsoft Collaboration, Industry Learning, and Technical Leadership

  • Collaborate across wider Microsoft , Safety, Security, Responsible AI, Identity and partner teams to create shared platform capabilities for risk detection, abuse prevention, evidence generation, and enforcement governance.
  • Evangelize a coherent Trust & Safety AI platform strategy across Microsoft, helping teams converge on shared architectures, reusable abstractions, common taxonomies, and consistent decisioning patterns.
  • Evangelize  & Learn from industry peers, and approved partner networks facing similar scale abuse patterns, and translate those learnings into practical platform improvements.

 



Qualifications
  • Bachelor’s Degree in Computer Science or related technical field AND 15+ years of professional software engineering experience, or equivalent practical experience.8+ years of senior technical leadership experience influencing engineers, technical leads, architects, applied scientists, or cross-functional engineering teams across complex platform or product areas.
  • Proven experience architecting and delivering large-scale production systems with meaningful reliability, scalability, latency, correctness, availability, security, and operational requirements.
  • Deep technical experience in one or more of the following platform areas: AI/ML systems, agentic systems, model-serving infrastructure, decisioning systems, distributed systems, data platforms, workflow platforms, risk platforms, security platforms, or Trust & Safety systems.
  • Experience building or technically leading production AI/ML systems, including exposure to modern AI patterns such as LLM-based workflows, retrieval-augmented generation, model orchestration, automated reasoning, human-in-the-loop systems, AI-assisted operational tooling, or agentic workflows.
  • Strong understanding of the engineering requirements for deploying AI or decision systems in production, including evaluation, observability, quality measurement, rollout safety, fallback behavior, latency/cost tradeoffs, drift detection, explainability, governance, and operational reliability.
  • Experience designing high-integrity systems where decisions must be auditable, reproducible, explainable, governed, and secure, especially when handling sensitive signals, advertiser impact, policy enforcement, risk decisions, or compliance-sensitive workflows.
  • Ability to drive clarity from ambiguity, define technical direction, create reusable platform abstractions, and influence execution across multiple teams without relying on direct management authority.
  • Strong written and verbal communication skills, including the ability to explain architecture, tradeoffs, risks, sequencing, and technical strategy to senior engineering, science, product, policy, security, and business leaders.

Preferred Qualification
The candidate is not expected to have deep experience in every area below. Broader coverage across these areas is preferred, especially where the candidate has demonstrated the ability to connect multiple domains into durable platform architecture and production systems.

  • Deep domain experience in Trust & Safety, Fraud, Abuse, Risk, Security, Ads Quality, Marketplace Integrity, Policy Enforcement, or advertiser protection systems.Experience with adversarial systems, including phishing, malware, cloaking, account takeover, payment abuse, fake identities, compromised advertisers, coordinated fraud, and policy evasion.
  • Experience building Deep Research Agents, investigation agents, reviewer-assist systems, retrieval-augmented generation systems, LLM-powered operational workflows, or AI systems that produce grounded evidence for human or automated decisions.
  • Expertise in heterogeneous inference platforms supporting LLMs, SLMs, wide & deep models, ensembles, graph models, classical ML models, heuristics, and rules engines.
  • Experience with entity intelligence, knowledge graphs, web crawling, domain reputation, business identity resolution, provenance, evidence extraction, or risk scoring.
  • Experience designing human-in-the-loop review systems, appeals workflows, audit platforms, policy reasoning systems, or enforcement governance mechanisms.
  • Experience with large-scale measurement systems for false positives, false negatives, model drift, agent quality, policy quality, reviewer quality, enforcement stability, business impact, and operational health.
  • Experience collaborating with Trust, Safety, Security, Privacy, Identity, Compliance, Legal, or Responsible AI teams across multiple products or platforms.
  • Experience evangelizing technical strategy across multiple teams, learning from industry peers, and helping establish shared standards, taxonomies, schemas, signal-quality measures, or platform patterns.
  • Experience working with industry partners, trusted abuse-prevention networks, threat-intelligence providers, domain-reputation providers, identity-verification providers, payment-risk partners, or ecosystem safety initiatives. 

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.




Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar jobs

Senior Software Engineer

Redmond, US
Software Engineering

Senior Software Engineer

Redmond, US
Software Engineering

Principal Applied Scientist

Redmond, US
Applied Sciences

Noida, India · Redmond, United States · Mountain View, United States · Bengaluru, India · Hyderabad, India

Principal Software Engineer

Location
Noida, India · Redmond, United States · Mountain View, United States · Bengaluru, India · Hyderabad, India
Job Number
200045102-en-3
City
Noida, Redmond, Mountain View, Bengaluru, Hyderabad
Team
Content Engineering
Country
India, United States
Discipline
Software Engineering

Overview

Are you passionate about architecting, building, and maintaining next-generation platforms for real-time data delivery that power Microsoft’s multi-billion-dollar advertising business? On our team, you’ll design and evolve complex systems, apply AI and next-gen technologies to solve modern engineering challenges, and collaborate with Ads and Bing teams to enable new scenarios. We’re a fast-paced, inclusive team that values innovation, learning, and impact. If you’re self-driven and excited to tackle deep technical problems at scale while shaping the future of intelligent, next-gen platforms, we’d love to work with you.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Responsibilities

  • Define and drive the technical vision and architecture for large-scale, complex systems, ensuring scalability, reliability, security, and cost efficiency.
  • Hands-on, contributing to core code, complex implementations, and production issue resolution.
  • Influence and align multiple teams and stakeholders across organizations through solid technical leadership.
  • Partner with Product Management and leadership to translate business goals into robust technical solutions.
  • Mentor engineers, raise the technical bar, and promote a culture of engineering excellence.
  • Use AI tools and techniques to enhance engineering workflows, automate processes, and unlock new capabilities.
  • Continuously refine data pipelines and system architecture to improve performance, reliability, and cost efficiency.
  • Collaborate with Ads and Bing teams to enable new scenarios, integrate shared infrastructure, and deliver unified solutions.
  • Provide production support by fixing bugs and resolving live-site issues to ensure system availability and reliability.

Qualifications

Required/Minimum Qualifications:

  • Bachelor’s Degree in Computer Science or related technical field AND 10+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
  • Solid SQL expertise, Kafka, Hadoop, Hive and product development experience including memory management, multithreading, and performance optimization, cloud expertise, AI.
  • Proven experience designing and operating distributed, cloud-based systems at scale.
  • Ability to influence without authority and drive alignment across multiple teams.
  • Solid expertise in system architecture, data structures, algorithms, and software design patterns.
  • Experience with microservices, data pipelines, or large data systems.
  • Experience with CI/CD pipelines, version control systems (e.g., Git), and build tools

Preferred Qualifications:

  • Experience with AdTech is a big plus.
  • Experience building high-throughput, low-latency, or mission-critical services.
  • Deep knowledge of Azure or other hyperscale cloud platforms.
  • Experience with data-intensive systems, machine learning platforms, or large marketplaces.
  • Track record of leading major architectural initiatives or platform transformations.
  • Solid written and verbal communication skills, including technical design documentation.
  • A collaborative team player with curiosity, a growth mindset, and a solid sense of responsibility
  • Exposure to AI/ML concepts and practical application of AI tools to solve modern engineering challenges.

Other Requirements: Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings:

  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter. 

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.


Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar jobs

Senior Software Engineer

Redmond, US
Software Engineering

Senior Software Engineer

Redmond, US
Software Engineering

Principal Applied Scientist

Redmond, US
Applied Sciences

Noida, India

Principal Architect

Location
Noida, India
Job Number
200038111-en-1
City
Noida
Team
Teams
Country
India
Discipline
Software Engineering

Overview

Core Services team in Microsoft Teams provides the foundational infrastructure, platform services, network, security, monitoring and governance to run planet scale distributed systems and microservices architecture that powers Microsoft Teams. Security and Reliability are at the heart of what this team aspires to do day in and day out.  

As a Principal Architect in Teams Core Services team you will be responsible for the architecture and design of the platform services, supporting Infra and Network Security, front door, routing, gateway, CDN, DNS and monitoring layers for the microservices powering Microsoft Teams. This opportunity will allow you to hone your skills in improving your security acumen working with experts in the space and become adept at troubleshooting and securing the network layer, and improve the reliability of such mission critical layers investing in active-active architectures at every possible level.   

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. 

Responsibilities

  • You will drive strategic improvements in architecture and design standards across the Teams services, while prioritizing development and implementation efforts. You’ll also lead by example with a hands-on approach to architecture. 
  • You develop, test, and implement end-to-end optimization of code. Additionally, you develop, review, and provide feedback on code, scripts, systems, and/or platforms. You analyze data from telemetry pipelines and monitors tools that detail operations metrics. You identify optimal uses for existing tools and/or models to identify contributing factors or points of failure that are affecting the efficiency of complex systems. 
  • You will lead by example and mentors’ others to produce extensible and maintainable code used across products. 
  • You will leverage subject-matter expertise of cross-product features with appropriate stakeholders (e.g., project managers) to drive multiple group projects plans, release plans, and work items. 
  • You will hold accountability as a Designated Responsible Individual (DRI), mentoring engineers across products/solutions, working on call to monitor system/product/service for degradation, downtime, or interruptions. 
  • You will proactively seek new knowledge and adapt to new trends, technical solutions, and patterns that will improve the availability, reliability, efficiency, observability, and performance of products while also driving consistency in monitoring and operations at scale and sharing knowledge with other engineers. 

Qualifications

Required Qualifications:

  • Bachelor’s Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python.
    • OR equivalent experience.
  • 2+ years of experience handling security issues for large scale cloud services, network infrastructures, native applications, web applications, distributed and database systems.
  • 2+ years of experience on areas like TCP/IP concepts, load balancing, CDN, ACL, routing, TLS, Certificate Lifecycle management, IP network analysis and performance and application issues using standard tools. 

Preferred Qualifications:

  • Master’s Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python.
    • OR Bachelor’s Degree in Computer Science or related technical field AND 15+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python.
    • OR equivalent experience.
  • Experience leading cross-functional collaboration across multiple engineering organizations with strong incident response and ability to drive technical decisions in ambiguous environments. 
  • Experience with Network security, Network troubleshooting, Cloud Security, Security Policy management and Certificate lifecycle management. 
  • Knowledge of Cloud Infrastructure services or ‘Infrastructure as a Service [IaaS]’ which delivers computer infrastructure, typically a platform virtualization environment as a service. 
  • Knowledge of automation technologies, leveraging AI for productivity improvements, methods, and processes used for quality and cost improvements. 
  • Experience managing horizontal initiatives/programs that span multiple teams/services.

 

 

#CAPIDC

#MicrosoftTeams

#AIForDevProductivity

#DistributedSystems

#EngineeringManagement

#SiteReliabilityEngineering

#NetworkSecurity 

#CloudSecurity 

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.


Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar jobs

Senior Software Engineer

Redmond, US
Software Engineering

Senior Software Engineer

Redmond, US
Software Engineering

Principal Applied Scientist

Redmond, US
Applied Sciences

Noida, India

Senior Software Engineer – Kubernetes & IAC

Location
Noida, India
Job Number
200047362-en-2
City
Noida
Team
Agent 365
Country
India
Discipline
Software Engineering

Overview

Microsoft is a company where passionate innovators come to collaborate, envision what can be and take their careers to levels they cannot achieve anywhere else. This is a world of more possibilities, more innovation, and more openness in a cloud-enabled world. The Business & Industry Copilots group is a rapidly growing organization that is responsible for the Microsoft Dynamics 365 suite of products, Power Apps, Power Automate, Dataverse, AI Builder, Microsoft Industry Solution and more. Microsoft is considered one of the leaders in Software as a Service in the world of business applications and this organization is at the heart of how business applications are designed and delivered.  

Responsibilities

We are looking for a Senior Software Engineer to join our team!

This is an incredible time to be part of our team and contribute to a highly strategic initiative for Microsoft. Power Platform brings together multiple products designed to empower customers in their digital transformation journey. In this role, you’ll help build scalable, reliable, secure, and compliant AI infrastructure that powers every product across Power Platform. You’ll work alongside a passionate team of engineers who thrive on solving complex challenges at scale while delivering exceptional quality. If you’re excited about driving innovation and shaping the future of AI-powered solutions, we’d love to have you on board.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Qualifications

Job Responsibilities: 

  • Design, build, and operate cloud-scale, multi-tenant infrastructure platforms with a strong focus on reliability, security, scalability, and operational excellence.
  • Lead the design and implementation of Kubernetes-based infrastructure, including cluster architecture, networking, storage, and workload isolation strategies.
  • Build and evolve Infrastructure as Code (IaC) solutions (e.g., ARM, Bicep, Terraform, Helm) to enable repeatable, auditable, and automated infrastructure provisioning and lifecycle management.
  • Apply systems thinking to design for failure domains, including regional isolation, availability zone strategies, dependency management, and blast-radius reduction.
  • Drive resiliency and reliability improvements through proactive design reviews, fault modeling, chaos testing, and post-incident learning.
  • Build tooling and automation to detect, diagnose, and self-heal infrastructure and platform issues, enabling customers and support teams to self-resolve problems.
  • Identify recurring operational issues and escalation patterns, and drive engineering solutions such as self-healing mechanisms, automation, guardrails, and platform abstractions.
  • Partner closely with product, SRE, and Azure platform teams to define regional deployment strategies, capacity planning, and safe rollout patterns.
  • Contribute to product and platform improvements by filing impactful bugs, proposing design changes, and shipping fixes to production to prevent customer impact.
  • Communicate complex technical issues and recommendations clearly and concisely, influencing cross-team decisions and driving measurable business outcomes.
  • Embody Microsoft’s culture and values, contributing to a collaborative, inclusive, and growth-oriented engineering environment.

Required Qualifications: 

  • Bachelor’s Degree in Computer Science or related technical field AND 8+ years of technical engineering experience
    OR equivalent experience.
  • Hands-on experience with Kubernetes or container orchestration platforms in production environments.
  • Experience with Infrastructure as Code (IaC) tools such as ARM, Bicep, Terraform, or similar.

Other Requirements: Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings:

  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter. 

Preferred Qualifications: 

  • Strong understanding of distributed systems concepts, including failure domains, consistency, availability, and fault tolerance.
  • Experience designing for high availability, regional resiliency, and disaster recovery in cloud environments.
  • Familiarity with cloud networking, storage, and security fundamentals, including identity, access control, and isolation boundaries.
  • Experience operating large-scale services with an emphasis on reliability engineering, incident response, and postmortem-driven improvements.
  • Ability to reason about blast radius, dependency management, and safe rollout strategies in complex systems.

#CAPJobs #Agent365jobs

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.


Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar jobs

Senior Software Engineer

Redmond, US
Software Engineering

Senior Software Engineer

Redmond, US
Software Engineering

Principal Applied Scientist

Redmond, US
Applied Sciences

Noida, India

Senior Software Engineer – Kubernetes & IAC

Location
Noida, India
Job Number
200047362-en-5
City
Noida
Team
Agent 365
Country
India
Discipline
Software Engineering

Overview

Microsoft is a company where passionate innovators come to collaborate, envision what can be and take their careers to levels they cannot achieve anywhere else. This is a world of more possibilities, more innovation, and more openness in a cloud-enabled world. The Business & Industry Copilots group is a rapidly growing organization that is responsible for the Microsoft Dynamics 365 suite of products, Power Apps, Power Automate, Dataverse, AI Builder, Microsoft Industry Solution and more. Microsoft is considered one of the leaders in Software as a Service in the world of business applications and this organization is at the heart of how business applications are designed and delivered.  

Responsibilities

We are looking for a Senior Software Engineer to join our team!

This is an incredible time to be part of our team and contribute to a highly strategic initiative for Microsoft. Power Platform brings together multiple products designed to empower customers in their digital transformation journey. In this role, you’ll help build scalable, reliable, secure, and compliant AI infrastructure that powers every product across Power Platform. You’ll work alongside a passionate team of engineers who thrive on solving complex challenges at scale while delivering exceptional quality. If you’re excited about driving innovation and shaping the future of AI-powered solutions, we’d love to have you on board.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Qualifications

Job Responsibilities: 

  • Design, build, and operate cloud-scale, multi-tenant infrastructure platforms with a strong focus on reliability, security, scalability, and operational excellence.
  • Lead the design and implementation of Kubernetes-based infrastructure, including cluster architecture, networking, storage, and workload isolation strategies.
  • Build and evolve Infrastructure as Code (IaC) solutions (e.g., ARM, Bicep, Terraform, Helm) to enable repeatable, auditable, and automated infrastructure provisioning and lifecycle management.
  • Apply systems thinking to design for failure domains, including regional isolation, availability zone strategies, dependency management, and blast-radius reduction.
  • Drive resiliency and reliability improvements through proactive design reviews, fault modeling, chaos testing, and post-incident learning.
  • Build tooling and automation to detect, diagnose, and self-heal infrastructure and platform issues, enabling customers and support teams to self-resolve problems.
  • Identify recurring operational issues and escalation patterns, and drive engineering solutions such as self-healing mechanisms, automation, guardrails, and platform abstractions.
  • Partner closely with product, SRE, and Azure platform teams to define regional deployment strategies, capacity planning, and safe rollout patterns.
  • Contribute to product and platform improvements by filing impactful bugs, proposing design changes, and shipping fixes to production to prevent customer impact.
  • Communicate complex technical issues and recommendations clearly and concisely, influencing cross-team decisions and driving measurable business outcomes.
  • Embody Microsoft’s culture and values, contributing to a collaborative, inclusive, and growth-oriented engineering environment.

Required Qualifications: 

  • Bachelor’s Degree in Computer Science or related technical field AND 8+ years of technical engineering experience
    OR equivalent experience.
  • Hands-on experience with Kubernetes or container orchestration platforms in production environments.
  • Experience with Infrastructure as Code (IaC) tools such as ARM, Bicep, Terraform, or similar.

Other Requirements: Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings:

  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter. 

Preferred Qualifications: 

  • Strong understanding of distributed systems concepts, including failure domains, consistency, availability, and fault tolerance.
  • Experience designing for high availability, regional resiliency, and disaster recovery in cloud environments.
  • Familiarity with cloud networking, storage, and security fundamentals, including identity, access control, and isolation boundaries.
  • Experience operating large-scale services with an emphasis on reliability engineering, incident response, and postmortem-driven improvements.
  • Ability to reason about blast radius, dependency management, and safe rollout strategies in complex systems.

#CAPJobs #Agent365jobs

This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.


Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar jobs

Senior Software Engineer

Redmond, US
Software Engineering

Senior Software Engineer

Redmond, US
Software Engineering

Principal Applied Scientist

Redmond, US
Applied Sciences
English (United States)
Your Privacy Choices Opt-Out Icon Your Privacy Choices
Consumer Health Privacy Sitemap Contact Microsoft Privacy Manage cookies Terms of use Trademarks Safety & eco Recycling About our ads