Director of Product Management, Agentforce Voice Models
About this role
Employer-provided description, formatted for easier reading.
About the Role
Agentforce is Salesforce's next-generation AI platform, delivering autonomous agents that reason, take action, and communicate naturally across every customer touchpoint. Voice is the fastest-growing interaction surface—from contact-center automation and field-service assistants to real-time sales coaching and multilingual global deployments.
As Director of Agentforce Voice Models, you will own the full voice intelligence stack: Automatic Speech Recognition (ASR / STT), Text-to-Speech (TTS), Speech-to-Speech (S2S) end-to-end pipelines, speaker diarization, prosody modeling, and the language-coverage roadmap that makes Agentforce sound natural in every market we serve.
You will partner with product, infrastructure, and go-to-market teams to set the bar for transcription accuracy, latency, and voice expressiveness at enterprise scale.
What You'll Do
Voice Model Strategy & Roadmap
- Define and own the multi-year roadmap for Agentforce voice capabilities, spanning ASR/STT, TTS, S2S, and real-time voice agents.
- Set accuracy, latency, and quality benchmarks (WER, MOS, RTF, DMOS) and drive the organization to meet them.
- Evaluate build vs. buy vs. partner decisions for new voice model capabilities and maintain relationships with key academic and industry partners.
ASR / STT (Automatic Speech Recognition)
- Lead the development of production-grade ASR systems optimized for telephony, WebRTC, and device-side deployment.Word Error Rate (WER) improvement across noise conditions, accents, and domain-specific vocabularyStreaming and batch recognition pipelines with sub-200 ms first-token latency targetsCustom vocabulary and language model adaptation (hot-word boosting, domain LM interpolation)Punctuation restoration and inverse text normalization (ITN) for downstream NLU
- Drive multilingual and code-switching ASR coverage across priority languages; govern the language onboarding process including data acquisition, model training, and acceptance testing.
TTS (Text-to-Speech)
- Own the neural TTS pipeline—voice cloning, persona design, SSML compliance, and real-time synthesis—for Agentforce agent personas.
- Lead prosody research: intonation, rhythm, stress, and pause modeling that produces natural-sounding enterprise voices across conversational contexts.
- Manage voice talent agreements, ethical AI review, and consent frameworks for synthetic voice creation.
- Drive naturalness, expressiveness, and brand-consistency quality bars using subjective (MOS, CMOS) and objective (mel-cepstral distortion) evaluation frameworks.
Speech-to-Speech (S2S) & Real-Time Voice Agents
- Architect low-latency S2S pipelines that enable full-duplex conversational AI without the ASR→NLU→TTS handoff penalty.
- Partner with the Agentforce Reasoning team to integrate voice understanding with agent action loops (tool calls, CRM lookups, escalation routing).
- Establish interruption, barge-in, and turn-taking models appropriate for enterprise voice agents.
Speaker Diarization & Voice Analytics
- Deliver production speaker diarization ("who spoke when") for multi-party calls, enabling accurate per-speaker transcripts used in call coaching, compliance, and analytics.
- Develop speaker verification and voice-biometric capabilities for secure agent authentication use cases.
- Partner with the Einstein Analytics team to surface voice-derived signals (sentiment, engagement, talk-time ratios) in Salesforce dashboards.
Language Coverage & Localization
- Own the global language support roadmap; prioritize languages by customer demand, addressable market, and data availability.
- Establish data governance, annotation, and quality-control pipelines for low-resource languages.
- Work with regional Salesforce teams to validate dialect, accent, and cultural appropriateness of voice personas.
Engineering Leadership & Team Building
- Recruit, develop, and retain a world-class team of research engineers, applied scientists, and ML engineers (target team size: 20–30).
- Set technical direction, drive architectural decisions, and maintain engineering excellence through code review culture, rigorous evaluation frameworks, and production incident reviews.
- Build a culture of experimentation: rapid A/B testing, red-teaming for voice safety, and continuous model refreshes.
Cross-Functional Partnership
- Partner with Product Management to translate customer and regulatory requirements into model specifications and acceptance criteria.
- Collaborate with Legal, Privacy, and Trust & Safety on voice AI ethics, speaker consent, deepfake-detection, and compliance (HIPAA, PCI, GDPR).
- Represent Agentforce Voice at external conferences, standards bodies (W3C, IETF), and customer briefings.
Who You Are
Required Experience & Skills
- 10+ years in speech/audio machine learning, with 4+ years in a senior leadership role managing teams of 10 or more engineers or scientists.
- Deep hands-on expertise in at least two of: ASR (end-to-end or hybrid CTC/attention architectures), neural TTS (VITS, Voicebox, Matcha-TTS, or equivalent), or real-time speech processing pipelines.
- Proven track record of shipping production voice models at scale (millions of minutes per day) with measurable accuracy and latency improvements.
- Fluency in prosody modeling: pitch contour modeling, duration prediction, and expressiveness fine-tuning.
- Experience with speaker diarization systems (clustering-based, end-to-end, or hybrid).
- Strong understanding of multilingual and multi-dialect NLP/ASR challenges; experience with at least 5 languages in production.
- Proficiency in Python; experience with PyTorch or JAX model training at scale; familiarity with cloud-based training infrastructure (AWS, GCP, or Azure).
- Excellent written and verbal communication—able to translate complex model behavior into product and business narratives for executives.
Preferred Qualifications
- PhD in Computer Science, Electrical Engineering, Linguistics, or related field with a speech/audio focus, or equivalent industry experience.
- Familiarity with streaming inference, model quantization (INT8/INT4), and on-device deployment for low-latency voice.
- Experience with voice safety, watermarking, and deepfake/spoofing detection.
- Background in telephony protocols (SIP, RTP, WebRTC) and enterprise contact-center platforms (Genesys, NICE, Avaya, Amazon Connect).
- Published research or patents in speech processing, audio ML, or a closely related field.
- Experience operating within regulated industries (financial services, healthcare) with voice data compliance requirements.
Why Salesforce & Agentforce
At Salesforce, we believe AI agents will redefine how businesses and people work. Agentforce is at the center of that shift—and Voice is its most human interface. You will:
- Work at the intersection of frontier research and enterprise-grade reliability, with direct customer impact across 150,000+ organizations worldwide.
- Have access to proprietary, enterprise-grade voice data, world-class compute, and a platform with distribution that no startup can match.
- Operate with the autonomy of a startup within the resources of a $35B+ company.
- Be part of Salesforce's commitment to Responsible AI: building technology that is accurate, fair, transparent, and safe for every voice we serve.