Machine Learning Engineer - Evals
About this role
Employer-provided description, formatted for easier reading.
Visa sponsorship details: Not open to any visas (e. g. US citizen, Green Card holders)
Salary : $220K - $300K
On-site work policy: 5 days in-office in New York, NY
Full-time position
Location: New York, NY
Equity: Competitive equity
About this role
You'll own evaluation of Honcho end to end, the instrument that tells a research-driven team whether their identity representations are actually getting better. We build the memory and identity layer for the agentic world, and the product works.
It's a complex multi-agent harness with real black-box behavior, so the job is defining what "better" means for representations that change over time, without ground truth, and building the machinery that measures it.
If you love constructing measurements from nothing, reading traces to find what's actually broken, and shipping the fix yourself, you'll feel at home here.
What You'll Do
- Design the evals.
- Define the scores and decide what "better" means for an entity representation that changes over time, and keep that definition current as the product and methods move.
- Build the pipelines and harnesses.
- Data in, labels, versions, reruns, judges: the machinery that lets the team ask a new question this week and get an answer this week.
- Run them and harvest insights.
- Read the traces and results, find what's actually broken, not what's easy to measure, and propose fixes that produce higher-fidelity representations.
- Build simulation agents at scale.
- The ceiling on iteration speed is how many entities can be modeled and measured at once.
- You raise it.
- Own the loop end to end.
- The question, the pipeline, the rerun, the writeup.
- Nothing gets scoped and handed off, and nothing waits on someone else's sprint.
Requirements
Work experience
- Must have
: Designed, built, and ran evals end to end (LLMS)
- Must have:
Production-quality Python (tests, maintenance, release cadence)
Seniority
- Required
: 3+ years of experience 3+ years designing and owning LLM evaluation
- Required
: PyTorch, Hugging Face, Weights & Biases. Data pipeline tooling (Hydra, Kafka, SQL) a plus
Work experience
- Nice to have:
Work on memory, identity, personalization, agent evaluation, or agentic optimization.
- Nice to have:
Shipped research: first or co-author at NeurIPS, ICML, ICLR, or equivalent, OR self-published open-source work of comparable quality
Education
- Nice to have:
Degree in computer science, mathematics, statistics, or a related quantitative field