Principal Applied Scientist

Microsoft · Redmond, WA, US

Spotted 2h ago

What you'll need to apply

Fields this application requires

RésuméNameEmailPhoneLocationWork authorization answerVisa sponsorship answer

Company-specific questions

  • Is your legal name the same as your preferred name?
  • Address
  • Country/region of residence
  • State
  • Postal Code/Zip
  • Ethnicity/Race
  • Gender
  • U.S. Armed Forces Status
  • Veteran Status
  • Are you currently or have you ever been a member of the military, a civilian employee, or an official of any government, whether national, state, local, or foreign?
  • Have you signed a non-compete or non-disclosure statement which may become an obstacle to your acceptance at Microsoft?
  • Have you ever worked with Microsoft as a full-time / part-time employee, intern, vendor, agency temporary, or business guest? If yes, please provide as much information about your former employment as you can.
  • Are you currently employed by a Microsoft subsidiary (e.g., LinkedIn, GitHub, Gaming Studios, Activision, Blizzard, King)?
  • As part of the online application process you were asked whether you possess certain minimum required qualifications for the role to which you are applying. By selecting yes, you agree that you answered these questions accurately. You further acknowledge that your answers may result in your application not being considered further for this role if you do not currently meet the required qualifications for the role.
  • By checking this you agree to the Microsoft Data Privacy Notice (DPN) .
  • By checking this, you affirm that you have familiarized yourself with the Microsoft recruiting process and agree to the candidate code of conduct .
Job description

About this role

Employer-provided description, formatted for easier reading.

Overview Microsoft Cloud Operations and Innovation (CO+I) underpins Microsoft’s global cloud infrastructure, driving the innovation, planning, design, construction, and operation of one of the largest data center fleets in the world. We are seeking an exceptional Principal Applied Scientist to join the Data Center Applied AI team.

In this role, you will play a pivotal role in advancing and integrating cutting‑edge, multimodal agentic AI systems into core tools and operational workflows that power Microsoft’s data centers. Your work will drive operational efficiency at scale and help advance Microsoft’s mission to empower every person and every organization on the planet to achieve more through intelligent, multimodal AI agents.

As a Principal Applied Scientist t, you will bridge state‑of‑the‑art AI research with production‑grade engineering , delivering agent architectures that are reliable, secure, scalable, observable, and measurable .

You will collaborate closely with Business and Engineering to build and deploy agentic capabilities across workflow automation, information retrieval, retrieval‑augmented generation (RAG), tool and function calling, long‑horizon task execution, and multi‑agent orchestration.

This role blends deep AI and applied science expertise with strong engineering judgment, operational rigor, and a bias for action. Success requires end‑to‑end ownership, a growth mindset, and strong customer empathy. Join us to shape the future of agentic AI for data center operations at global scale.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. We work with a growth mindset, innovate to empower others, and collaborate to realize shared goals—guided by our values of respect, integrity, and accountability to foster an inclusive culture where everyone can thrive.

Responsibilities

  • Advance Applied AI Research to Improve Quality and Reliability
  • Apply deep expertise in Generative AI, deep learning, NLP, and multimodal models to translate cutting-edge research into high-impact, production-ready AI solutions .
  • Design and execute experiments that measurably improve agent planning, memory, grounding, reasoning, and long‑horizon task completion .
  • Perform lightweight fine‑tuning (e.g., LoRA and related techniques) on multimodal and large multimodal language models—to improve entity recognition, reasoning accuracy, and response quality.
  • Implement state‑of‑the‑art approaches using foundation models, advanced prompt engineering, RAG, knowledge graphs, and multi‑agent architectures , complemented by classical ML techniques where appropriate.
  • Build and Ship Scalable Agentic AI Systems
  • Architect and implement end‑to‑end agent workflows that decompose user intent into executable plans, intelligently select and invoke tools, and recover gracefully from errors and edge cases.
  • Rapidly prototype solutions and partner with engineering teams to drive production deployment , including debugging live systems and building AIOps workflows .
  • Use data and telemetry to identify AI quality gaps, generate insights, and deliver proofs of concept that apply research innovations to real‑world operational challenges.
  • Design robust multi‑step reasoning and tool‑use strategies , including function calling, code execution, APIs, and secure connectors, with strong safety and reliability guardrails.
  • Drive production excellence across latency, reliability, cost efficiency, observability, monitoring, and safe fallback behaviors .
  • Own Evaluation, Metrics, and Iteration Loops
  • Define task‑level evaluation metrics for agentic behavior, including success rate, tool‑call accuracy, step efficiency, hallucination rate, safety violations, time‑to‑completion, and user satisfaction .
  • Build and maintain offline and online evaluation pipelines , including:
  • Golden datasets, scenario simulators, and regression test suites
  • Human‑in‑the‑loop evaluation frameworks and rubric design
  • A/B experimentation and telemetry‑driven iteration loops
  • Collaborate Across Disciplines and Lead Technical Execution
  • Partner with business and engineering stakeholders to translate requirements into clear technical specifications, research plans, and delivery milestones .
  • Lead design reviews and influence engineering decisions across agent frameworks, model integration, and system architecture .
  • Mentor and support other applied scientists and engineers, raising the bar for applied AI execution and production quality

Qualifications

Required/minimum qualifications

  • Bachelor's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 8+ years related experience (e.g., statistics, predictive analytics, research)
  • OR Master's Degree in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 6+ years related experience (e.g., statistics, predictive analytics, research)
  • OR Doctorate in Statistics, Econometrics, Computer Science, Electrical or Computer Engineering, or related field AND 5+ years related experience (e.g., statistics, predictive analytics, research)
  • OR equivalent experience

Other Requirements:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings: Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.

Preferred Qualifications

  • PhD in Computer Science, Machine Learning, AI, NLP, or related field , OR MS with significant experience delivering production‑grade AI systems .
  • 10+ years of hands‑on experience applying NLP, and/or LLMs to real‑world production problems .
  • Strong foundation in Large Language Models (LLMs) , Generative AI, deep learning, and modern NLP.
  • Proven experience building production‑grade LLM systems , including prompt engineering, RAG, and multi‑step reasoning pipelines.
  • Solid expertise in chunking strategies and embedding selection , including tradeoffs in retrieval quality, context size, and downstream task performance.
  • Experience designing and operating vector databases , including indexing strategies, similarity metrics, and retrieval optimization.
  • Experience supporting production operations (AIOps/MLOps) , including monitoring, logging, quality metrics, and debugging live systems.
  • Proficiency in Python and experience with modern ML frameworks (e.g., PyTorch, TensorFlow) and model serving/inference.
  • Strong applied experimentation skills: defining metrics, analyzing results, and iterating based on data.
  • Ability to collaborate effectively with engineering, product, and business stakeholders.
  • Experience with Cloud technology Stack
  • Experience with parameter‑efficient fine‑tuning techniques (e.g., LoRA) for LLMs or multimodal models.
  • Hands‑on experience building agentic or tool‑augmented LLM systems , including function calling, planners, and API/tool integration.
  • Experience with advanced RAG architectures , such as hybrid retrieval, reranking, grounding, and citation strategies.
  • Strong AIOps depth , including quality drift detection, alerting, rollback, and telemetry‑driven optimization.
  • Experience optimizing systems for latency, reliability, scalability, and cost efficiency at enterprise scale.
  • Prior experience working on mission‑critical or large‑scale enterprise AI systems .
  • Demonstrated mentorship or technical leadership within applied science or engineering teams.

#COICareers | #EPCCareers | #DCDCareers

Applied Sciences IC6 - The typical base pay range for this role across the U.S. is USD $165,600 - $296,400 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in

Interested in this role?Continue on Microsoft's careers page.
Apply on Microsoft