Director of Product Management, Agentforce Voice Models

Salesforce · California - San Francisco · Bellevue · San Francisco

Spotted 1h agoFull time
Job description

About this role

Employer-provided description, formatted for easier reading.

About the Role

Agentforce is Salesforce's next-generation AI platform, delivering autonomous agents that reason, take action, and communicate naturally across every customer touchpoint. Voice is the fastest-growing interaction surface—from contact-center automation and field-service assistants to real-time sales coaching and multilingual global deployments.

As Director of Agentforce Voice Models, you will own the full voice intelligence stack: Automatic Speech Recognition (ASR / STT), Text-to-Speech (TTS), Speech-to-Speech (S2S) end-to-end pipelines, speaker diarization, prosody modeling, and the language-coverage roadmap that makes Agentforce sound natural in every market we serve.

You will partner with product, infrastructure, and go-to-market teams to set the bar for transcription accuracy, latency, and voice expressiveness at enterprise scale.

What You'll Do

Voice Model Strategy & Roadmap

  • Define and own the multi-year roadmap for Agentforce voice capabilities, spanning ASR/STT, TTS, S2S, and real-time voice agents.
  • Set accuracy, latency, and quality benchmarks (WER, MOS, RTF, DMOS) and drive the organization to meet them.
  • Evaluate build vs. buy vs. partner decisions for new voice model capabilities and maintain relationships with key academic and industry partners.

ASR / STT (Automatic Speech Recognition)

  • Lead the development of production-grade ASR systems optimized for telephony, WebRTC, and device-side deployment.Word Error Rate (WER) improvement across noise conditions, accents, and domain-specific vocabularyStreaming and batch recognition pipelines with sub-200 ms first-token latency targetsCustom vocabulary and language model adaptation (hot-word boosting, domain LM interpolation)Punctuation restoration and inverse text normalization (ITN) for downstream NLU
  • Drive multilingual and code-switching ASR coverage across priority languages; govern the language onboarding process including data acquisition, model training, and acceptance testing.

TTS (Text-to-Speech)

  • Own the neural TTS pipeline—voice cloning, persona design, SSML compliance, and real-time synthesis—for Agentforce agent personas.
  • Lead prosody research: intonation, rhythm, stress, and pause modeling that produces natural-sounding enterprise voices across conversational contexts.
  • Manage voice talent agreements, ethical AI review, and consent frameworks for synthetic voice creation.
  • Drive naturalness, expressiveness, and brand-consistency quality bars using subjective (MOS, CMOS) and objective (mel-cepstral distortion) evaluation frameworks.

Speech-to-Speech (S2S) & Real-Time Voice Agents

  • Architect low-latency S2S pipelines that enable full-duplex conversational AI without the ASR→NLU→TTS handoff penalty.
  • Partner with the Agentforce Reasoning team to integrate voice understanding with agent action loops (tool calls, CRM lookups, escalation routing).
  • Establish interruption, barge-in, and turn-taking models appropriate for enterprise voice agents.

Speaker Diarization & Voice Analytics

  • Deliver production speaker diarization ("who spoke when") for multi-party calls, enabling accurate per-speaker transcripts used in call coaching, compliance, and analytics.
  • Develop speaker verification and voice-biometric capabilities for secure agent authentication use cases.
  • Partner with the Einstein Analytics team to surface voice-derived signals (sentiment, engagement, talk-time ratios) in Salesforce dashboards.

Language Coverage & Localization

  • Own the global language support roadmap; prioritize languages by customer demand, addressable market, and data availability.
  • Establish data governance, annotation, and quality-control pipelines for low-resource languages.
  • Work with regional Salesforce teams to validate dialect, accent, and cultural appropriateness of voice personas.

Engineering Leadership & Team Building

  • Recruit, develop, and retain a world-class team of research engineers, applied scientists, and ML engineers (target team size: 20–30).
  • Set technical direction, drive architectural decisions, and maintain engineering excellence through code review culture, rigorous evaluation frameworks, and production incident reviews.
  • Build a culture of experimentation: rapid A/B testing, red-teaming for voice safety, and continuous model refreshes.

Cross-Functional Partnership

  • Partner with Product Management to translate customer and regulatory requirements into model specifications and acceptance criteria.
  • Collaborate with Legal, Privacy, and Trust & Safety on voice AI ethics, speaker consent, deepfake-detection, and compliance (HIPAA, PCI, GDPR).
  • Represent Agentforce Voice at external conferences, standards bodies (W3C, IETF), and customer briefings.

Who You Are

Required Experience & Skills

  • 10+ years in speech/audio machine learning, with 4+ years in a senior leadership role managing teams of 10 or more engineers or scientists.
  • Deep hands-on expertise in at least two of: ASR (end-to-end or hybrid CTC/attention architectures), neural TTS (VITS, Voicebox, Matcha-TTS, or equivalent), or real-time speech processing pipelines.
  • Proven track record of shipping production voice models at scale (millions of minutes per day) with measurable accuracy and latency improvements.
  • Fluency in prosody modeling: pitch contour modeling, duration prediction, and expressiveness fine-tuning.
  • Experience with speaker diarization systems (clustering-based, end-to-end, or hybrid).
  • Strong understanding of multilingual and multi-dialect NLP/ASR challenges; experience with at least 5 languages in production.
  • Proficiency in Python; experience with PyTorch or JAX model training at scale; familiarity with cloud-based training infrastructure (AWS, GCP, or Azure).
  • Excellent written and verbal communication—able to translate complex model behavior into product and business narratives for executives.

Preferred Qualifications

  • PhD in Computer Science, Electrical Engineering, Linguistics, or related field with a speech/audio focus, or equivalent industry experience.
  • Familiarity with streaming inference, model quantization (INT8/INT4), and on-device deployment for low-latency voice.
  • Experience with voice safety, watermarking, and deepfake/spoofing detection.
  • Background in telephony protocols (SIP, RTP, WebRTC) and enterprise contact-center platforms (Genesys, NICE, Avaya, Amazon Connect).
  • Published research or patents in speech processing, audio ML, or a closely related field.
  • Experience operating within regulated industries (financial services, healthcare) with voice data compliance requirements.

Why Salesforce & Agentforce

At Salesforce, we believe AI agents will redefine how businesses and people work. Agentforce is at the center of that shift—and Voice is its most human interface. You will:

  • Work at the intersection of frontier research and enterprise-grade reliability, with direct customer impact across 150,000+ organizations worldwide.
  • Have access to proprietary, enterprise-grade voice data, world-class compute, and a platform with distribution that no startup can match.
  • Operate with the autonomy of a startup within the resources of a $35B+ company.
  • Be part of Salesforce's commitment to Responsible AI: building technology that is accurate, fair, transparent, and safe for every voice we serve.
Interested in this role?Continue on Salesforce's careers page.
Apply on Salesforce