Staff Software Engineer
What you'll need to apply
What this employer's standard application typically asks
About this role
Employer-provided description, formatted for easier reading.
This role is for a staff engineer who independently leads the design, development, and delivery of complex, high-impact AI-native and cloud-native systems. The IC4 engineer takes ownership of critical components or features, sets technical direction within their domain, and serves as a technical resource to peers and junior engineers.
They combine deep software engineering fundamentals with hands-on expertise in AI/ML integration, cloud architecture, and production systems. This role requires the ability to balance technical depth with organizational impact, working across teams to solve ambiguous problems and drive platform improvements.
Responsibility Area
Technical Leadership & Architecture
- Lead the design and architecture of complex, production-critical AI-native systems, including LLM integrations, agentic workflows, retrieval pipelines, and inference services
- Own end-to-end delivery of medium-to-large features or systems, including specification, design, implementation, testing, deployment, and production monitoring
- Make informed architectural and technical trade-off decisions between latency, cost, scalability, reliability, and model quality
- Establish evaluation frameworks, testing strategies, and production monitoring for AI-driven features; drive adoption of eval-driven development practices
- Design and optimize high-performance data pipelines, embedding systems, and vector stores handling millions of rows and complex retrieval scenarios
- Champion security best practices, including prompt-injection defense, data privacy, PII handling, responsible AI guardrails, and compliance with regulatory requirements
- Participate in architectural reviews across teams and contribute to platform-level technical decisions
- Stay current with emerging AI/ML frameworks, services, and techniques; evaluate and recommend adoption where applicable
Responsibility Area: Mentorship & Technical Influence
- Mentor junior and mid-level engineers (IC2/IC3) through pair programming, design reviews, architecture discussions, and hands-on coaching
- Conduct thorough code reviews, including evaluation of prompts, agent logic, retrieval strategies, and AI/model behavior validation
- Identify technical debt and improvement opportunities; advocate for and lead initiatives to address them
- Share knowledge through documentation, brown-bag sessions, and architectural guidance
- Contribute to hiring and interview processes; provide technical assessment and feedback on engineering candidates
- Model best practices in software craftsmanship, testing discipline, and debugging methodology
- Support onboarding of new team members and ensure knowledge transfer across projects
Responsibility Area: Cross-Functional Collaboration & Communication
- Work closely with product management, domain specialists, and other engineering teams to define requirements and technical specifications
- Translate complex technical concepts and AI trade-offs for non-technical stakeholders (product, leadership, customers)
- Lead root cause analysis for customer-reported issues, including those involving model behavior, hallucinations, retrieval quality, and distributed system complexity
- Participate in design reviews, technical planning sessions, and retrospectives; contribute perspectives on technical feasibility and risk
- Communicate clearly about design decisions, implementation challenges, AI limitations, and production incidents
- Build and maintain relationships with adjacent teams and external partners where applicable
- Proactively identify opportunities to improve development processes, tooling, and team efficiency
Responsibility Area: Software Development Excellence
- Develop robust, maintainable, and well-tested AI-native software using Python, Java, and JavaScript as appropriate
- Build comprehensive unit tests, integration tests, and AI evaluation harnesses (accuracy, regression, hallucination detection)
- Troubleshoot complex issues in distributed systems, probabilistic models, and cloud infrastructure; isolate root causes efficiently
- Implement secure coding practices and responsible AI guardrails throughout the development lifecycle
- Manage deployment, monitoring, and iteration of features in production; respond to performance issues and model drift
- Work with cloud services (AWS/GCP/Azure), managed AI services (Bedrock, Vertex AI, Azure OpenAI), Kubernetes, and CI/CD pipelines
- Contribute to code and architecture standards; ensure adherence to established best practices
Responsibility Area: ServiceNow Platform Development
- Design and develop custom applications, extensions, and integrations on the ServiceNow platform
- Develop workflows, business rules, and automation using ServiceNow scripting (JavaScript, GlideScript, REST APIs)
- Build data models, forms, and dashboards for ServiceNow applications; optimize database queries for performance
- Integrate third-party AI/ML services and custom APIs with ServiceNow platform services
- Contribute to and maintain ServiceNow plugin and app architecture; follow ServiceNow best practices and coding standards
- Participate in ServiceNow upgrades and configuration management; ensure compatibility with platform updates
- Mentor junior engineers on ServiceNow platform capabilities and development patterns
Responsibility Area: Customer Support & Product Excellence
- Participate in customer issue triage and root cause analysis for production ServiceNow application issues
- Work closely with customer success and support teams to understand and resolve complex technical issues
- Provide technical guidance to customers on application usage, configuration, and troubleshooting
- Identify patterns in customer issues and drive improvements to prevent recurrence
- Contribute to customer documentation and internal knowledge bases
- Support customer escalations; provide timely resolutions to critical issues affecting customer operations
- Gather customer feedback and communicate product improvement opportunities back to product management
- 10-14 years of practical software development experience, with 5+ years working with modern cloud-native and AI-native systems
- 3+ years of hands-on experience with LLM prompt engineering, agent design, and retrieval systems (RAG, semantic search, knowledge bases)
- 3+ years of demonstrated experience integrating third-party AI/ML services and platform APIs (OpenAI, Anthropic, AWS Bedrock, Google Vertex AI, etc.)
- 2+ years of hands-on experience developing applications and extensions on the ServiceNow platform
- Strong proficiency in JavaScript/GlideScript, Python, and Java; familiarity with ServiceNow scripting patterns and APIs
- Solid understanding of ServiceNow data model, form/workflow design, and business process automation
- Solid understanding of microservices architecture, REST/streaming APIs, and modern backend frameworks (Spring Boot, FastAPI, etc.)
- Experience designing and optimizing high-performance data pipelines, embedding systems, and vector stores at scale
- Demonstrated expertise in LLM orchestration frameworks (LangChain, LlamaIndex, Semantic Kernel) and vector databases (Pinecone, Weaviate, pgvector, etc.)
- Experience with cloud computing platforms and managed AI services (AWS, GCP, Azure); hands-on experience with infrastructure-as-code
- Proficiency with Kubernetes, containerization, CI/CD pipelines, and automated testing frameworks
- Experience developing and debugging evaluation harnesses, benchmarks, and metrics for AI/ML features
- Proven ability to troubleshoot and isolate root causes in complex, distributed, and probabilistic systems
- Strong communication skills; ability to explain technical concepts to both technical and non-technical audiences;