User Research

Technical Lead, Evaluation Infrastructure

Lead evaluation infrastructure that turns Nuro's autonomy logs and simulations into release-ready behavior evidence.

What the role actually is

Nuro is hiring a technical lead for the infrastructure behind autonomy evaluation and driverless safety validation. The team builds metrics frameworks, evaluation pipelines, analysis products, and tooling that turn simulation and road logs into evidence for L4 autonomy releases.

This is not a classic UX research job, but it belongs on the board because it is about measuring robot behavior in the world: whether the autonomous driver is ready, why it failed, and what evidence should gate deployment.

What you would work on

  • Lead technical direction for Nuro's autonomy evaluation infrastructure
  • Build pipelines and metrics that evaluate simulation and on-road autonomy behavior
  • Improve tooling that helps autonomy and safety teams understand model performance
  • Shorten the feedback loop between evaluation results, debugging, and release decisions

What they are asking for

  • Senior technical leadership in autonomy, robotics, ML infrastructure, simulation, or evaluation systems
  • Experience building reliable data or analysis platforms for complex technical teams
  • Ability to define metrics that explain behavior quality and safety readiness
  • Comfort partnering with autonomy, simulation, AI platform, and systems safety groups

Why this one is worth a look

Evaluation is becoming one of the most important disciplines in physical AI. Nuro's role is a good example: the job is less about producing a demo and more about building the evidence layer that decides whether autonomy is improving enough to ship.

About Nuro

Mountain View autonomy company developing AI-first driving systems for delivery and mobility platforms.

Visit Nuro

Related roles