Architect / Principal Software Engineer
About this role
Employer-provided description, formatted for easier reading.
We are looking for an Architect / Principal Software Engineer to lead the technical direction of our AI Engineering pods. You will split your time between hands-on engineering (about 30% to 40% prototyping core workflows, building reference implementations, and reviewing code) and higher-level system architecture, technical standards, and mentorship across teams.
You will work closely with US-based Product, Security, Data, and Infrastructure teams to make sure our agentic systems are reliable, fast, and secure. Your main focus will be designing predictable multi-agent workflows, low-latency LLM orchestration services, and retrieval pipelines that run cleanly in production.
Req. #1098579355
Responsibilities
- Set architectural direction across multiple AI pods, making clear trade-offs between model capabilities, cost, latency, and operational complexity
- Spend roughly a third of your time in code: writing backend AI services, building orchestration patterns, and defining integration contracts with frontend teams
- Design agent workflows with streaming responses, predictable error handling, and low latency for chat and automated tasks
- Establish shared standards for Model Context Protocol (MCP) servers, tool schemas, and context management
- Work with Okta Security and Infrastructure to enforce least-privilege tool execution, tenant isolation, and secret handling within agent loops
- Define observability requirements and service-level objectives, including tracing with tools like LangSmith or Arize Phoenix
- Mentor Staff and Senior engineers through direct design reviews and technical guidance
Requirements
- 10+ years of software engineering experience, with a background in designing, scaling, and operating distributed backend systems
- 6+ years of Python experience, especially with FastAPI, AsyncIO, and modern typing practices. Python is our primary language for backend and AI services
- 3+ years building production systems with LLMs, including LangChain, LangGraph, AWS Bedrock, Anthropic Claude, or OpenAI APIs. Candidates with classical ML backgrounds (PyTorch, scikit-learn) who moved into generative AI architectures are welcome
- Hands-on experience with LangGraph or similar stateful agent frameworks, particularly managing state graphs, loops, checkpointing, and human-in-the-loop steps
- 3 to 4 years of React and TypeScript experience, enough to review UI code, understand frontend state, and ensure streaming APIs work cleanly with the web client
- Production AWS experience, including EKS or ECS, Lambda, Bedrock, S3, Docker, and Terraform or OpenTofu
- Authentication and security fundamentals, specifically OAuth 2.0, OIDC, and secure credential handling
Nice to have
- Practical experience building custom MCP servers and defining tool discovery patterns
- Experience migrating from older retrieval systems (such as Amazon Kendra) to vector platforms like OpenSearch, Pinecone, or Qdrant
- Experience handling multimodal data (images, documents) within LangGraph state
- Background in Identity and Access Management (IAM) or Customer Identity (CIAM)
We offer
- Medical, Dental and Vision Insurance (Subsidized)
- Health Savings Account
- Flexible Spending Accounts (Healthcare, Dependent Care, Commuter)
- Short-Term and Long-Term Disability (Company Provided)
- Life and AD&D Insurance (Company Provided)
- Employee Assistance Program
- Unlimited access to LinkedIn learning solutions
- Matched 401(k) Retirement Savings Plan
- Paid Time Off – the employee will be eligible to accrue 15-25 paid days, depending on specific level and tenure with EPAM (accrual eligibility may change over time)
- Paid Holidays - nine (9) total per year
- Legal Plan and Identity Theft Protection
- Accident Insurance
- Employee Discounts
- Pet Insurance
- Employee Stock Purchase Program
- If otherwise eligible, participation in the discretionary annual bonus program
- If otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program
This Remote Position Cannot be Performed in New York City.
EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our clients, our employees, and our communities. We embrace a dynamic and inclusive culture.
Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
Engineer the Future with a Career at EPAM
This posting includes a good faith range of the salary EPAM would reasonably expect to pay the selected candidate. The range provided reflects base salary only. Individual compensation offers within the range are based on a variety of factors, including, but not limited to: geographic location, experience, credentials, education, training; the demand for the role; and overall business and labor market considerations.
Most candidates are hired at a salary within the range disclosed. Salary range: $140,000 - $165,000. In addition, the details highlighted in this job posting above are a general description of all other expected benefits and compensation for the position.
Applications will be accepted on a rolling basis.
In accordance with the LA County Fair Chance Ordinance, you may find a copy of the Notice containing a summary of the Ordinance’s key provisions here: Concept FCO Posting 8 27 24 (lacounty. gov)
EPAM will not provide new H-1B visa sponsorship for this position. Candidates with existing transferable H-1B status may be considered.
It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.