Back to jobs

Machine Learning Engineer – Evals

Mango Global Technologies · New York, NY, US

Spotted 3d agoFull-time

Job details

Pay
$220,000 – $300,000 a year
Work mode
On-site
Employment
Full-time
Level
Mid level
Experience
3+ years
Posted
Oct 8, 2026
Last confirmed open
Oct 8, 2026
Job description

About this role

Machine Learning Engineer – Evals

Location: New York, NY

Work arrangement: On-site, 5 days/week

Employment: Full-time

Compensation: $220,000–$300,000 base + competitive equity

Openings: 1

Visa sponsorship: Not available — candidates must already have unrestricted authorization to work in the United States.

About the Role

We are looking for a Machine Learning Engineer – Evals to own the evaluation of a sophisticated AI memory and identity system end to end.

You will define how we determine whether evolving entity representations are actually improving, build the infrastructure required to measure those improvements, and use evaluation results to identify and fix failures.

This is not a role for someone who simply runs established benchmarks. You will be expected to design novel evaluations from first principles, work with imperfect or absent ground truth, analyze model and agent traces, and turn research questions into production-quality evaluation systems.

The ideal candidate combines strong ML engineering skills with genuine research ability and enjoys working in an ambiguous, fast-moving environment.

What You’ll Do

  • Design evaluation systems Define what “better” means for representations that evolve over time.
  • Design novel evaluation methodologies where traditional ground truth is unavailable or insufficient.
  • Develop scoring frameworks, metrics, rubrics, judge models, and diagnostic methodologies.
  • Continuously evolve evaluation definitions as models and product capabilities change.
  • Build evaluation infrastructure Build production-quality Python pipelines and evaluation harnesses.
  • Manage datasets, labels, experiment versions, model versions, judges, reruns, and regression testing.
  • Create infrastructure that allows researchers and engineers to answer new evaluation questions quickly.
  • Integrate experiment tracking and observability into evaluation workflows.
  • Analyze results and improve the system Analyze model outputs, agent trajectories, and execution traces.
  • Identify meaningful failure modes rather than optimizing for superficial metrics.
  • Translate evaluation findings into concrete engineering or research improvements.
  • Run experiments, compare results, document findings, and drive the resulting fixes.
  • Build simulation agents at scale Develop and operate simulation environments capable of modeling large numbers of entities.
  • Increase the throughput and reliability of evaluation runs.
  • Build systems that enable rapid iteration across models, prompts, representations, and evaluation strategies.
  • Own the full evaluation loop You will own the process from: Research question → evaluation design → data → implementation → execution → analysis → diagnosis → rerun → recommendation There is no expectation that this work will be handed off between multiple teams.

Required Qualifications

Dealbreakers

3+ years of professional experience designing and owning LLM evaluations.

Demonstrated experience designing, building, and running LLM evaluations end to end.

Production-quality Python experience, including testing, maintainability, debugging, and release practices.

Experience creating evaluation methodologies rather than exclusively running existing benchmarks.

Strong ability to investigate model/agent failures using traces, outputs, datasets, and experimental results.

Ability to work on-site in New York City five days per week.

Must already have unrestricted U.S. work authorization; visa sponsorship is not available.

Required Technical Skills

Python

PyTorch

Hugging Face / Transformers

LLM evaluation and benchmarking

Experiment design and statistical/quantitative analysis

Evaluation datasets and labeling methodologies

Model/LLM-as-a-judge evaluation

ML experimentation and debugging

Preferred Technical Skills

Weights & Biases

Hydra

SQL

Kafka or other streaming/event-driven infrastructure

Evaluation/observability frameworks

Agent simulation and multi-agent systems

Large-scale data pipelines

Nice to Have

Experience with AI memory, identity, personalization, agent evaluation, or agentic optimization.

Research experience in LLMs, NLP, machine learning, evaluation, or related areas.

First-author/co-author research at NeurIPS, ICML, ICLR, ACL, EMNLP, or comparable venues.

Alternatively, a high-quality open-source project, benchmark, framework, or technical publication demonstrating comparable research ability.

Degree in Computer Science, Mathematics, Statistics, Machine Learning, or another quantitative discipline.

Who Will Not Be a Strong Fit

Candidates who have only run established benchmarks without designing their own evaluations.

Engineers whose LLM experience is limited to application development without evaluation ownership.

Candidates without substantial hands-on Python engineering experience.

Candidates without evidence of research thinking, novel experimentation, published work, or comparable open-source work.

Candidates unwilling or unable to work five days per week from the New York City office.

Compensation

$220,000–$300,000 base salary + competitive equity

Final compensation will depend on experience, technical depth, research contributions, and interview performance.

Hiring Process

The successful candidate should expect technical discussions focused on:

Designing an evaluation from scratch.

Diagnosing an LLM/agent failure from traces.

Building a production evaluation pipeline.

Research methodology and experimental design.

Python/ML engineering.

Previous research or open-source work.

Pay: $220,000.00 - $300,000.00 per year

Benefits:

  • Employee assistance program

Application Question(s):

  • Can you work full-time on-site in New York City 5 days per week, and do you currently have unrestricted U.S. work authorization without requiring visa sponsorship?
  • Have you either (a) published ML/LLM research at a recognized conference/journal or (b) released substantial open-source research/ML work of comparable quality?
  • Do you have hands-on experience with PyTorch AND Hugging Face/Transformers?
  • Do you have production-level Python experience, including writing tests, maintaining production code, debugging, and managing releases?
  • How many years of professional experience do you have designing and owning LLM evaluations?
  • Have you personally designed, built, and run an LLM evaluation system end to end, including defining the evaluation methodology/metrics rather than only executing an existing benchmark?
  • Share your Linkedin url

Work Location: In person

Interested in this role?Continue on Indeed to apply.
Apply on Indeed