Machine Learning Engineer - Evals

Ronaj Recruitment · New York, NY

Spotted 3h agofulltime
Job description

About this role

Employer-provided description, formatted for easier reading.

Visa sponsorship details: Not open to any visas (e. g. US citizen, Green Card holders)

Salary : $220K - $300K

On-site work policy: 5 days in-office in New York, NY

Full-time position

Location: New York, NY

Equity: Competitive equity

About this role

You'll own evaluation of Honcho end to end, the instrument that tells a research-driven team whether their identity representations are actually getting better. We build the memory and identity layer for the agentic world, and the product works.

It's a complex multi-agent harness with real black-box behavior, so the job is defining what "better" means for representations that change over time, without ground truth, and building the machinery that measures it.

If you love constructing measurements from nothing, reading traces to find what's actually broken, and shipping the fix yourself, you'll feel at home here.

What You'll Do

  • Design the evals.
  • Define the scores and decide what "better" means for an entity representation that changes over time, and keep that definition current as the product and methods move.
  • Build the pipelines and harnesses.
  • Data in, labels, versions, reruns, judges: the machinery that lets the team ask a new question this week and get an answer this week.
  • Run them and harvest insights.
  • Read the traces and results, find what's actually broken, not what's easy to measure, and propose fixes that produce higher-fidelity representations.
  • Build simulation agents at scale.
  • The ceiling on iteration speed is how many entities can be modeled and measured at once.
  • You raise it.
  • Own the loop end to end.
  • The question, the pipeline, the rerun, the writeup.
  • Nothing gets scoped and handed off, and nothing waits on someone else's sprint.

Requirements

Work experience

  • Must have

: Designed,​ built,​ and ran evals end to end (LLMS)

  • Must have:

Production-​quality Python (tests,​ maintenance,​ release cadence)

Seniority

  • Required

: 3+ years of experience 3+ years designing and owning LLM evaluation

  • Required

: PyTorch,​ Hugging Face,​ Weights &​ Biases. ​ Data pipeline tooling (Hydra,​ Kafka,​ SQL) a plus

Work experience

  • Nice to have:

Work on memory,​ identity,​ personalization,​ agent evaluation,​ or agentic optimization. ​

  • Nice to have:

Shipped research:​ first or co-​author at NeurIPS,​ ICML,​ ICLR,​ or equivalent,​ OR self-​published open-​source work of comparable quality

Education

  • Nice to have:

Degree in computer science,​ mathematics,​ statistics,​ or a related quantitative field

Interested in this role?Continue on Ronaj Recruitment's careers page.
Apply on Ronaj Recruitment