Applied Research Engineer

Location
San Francisco; New York
Employment type
Full-time
Location type
On-site
Department
Applied research

About the role

Applied research at Verona begins with a demanding real-world question: how do we know an AI system is useful, reliable, and improving on the workflow that matters? You will develop the evaluation methods, datasets, experiments, and learning systems that make those answers rigorous.

This role sits between research and production engineering. You will study failures from live enterprise deployments, form precise hypotheses about model and system behavior, build the infrastructure needed to test them, and turn the results into systems that improve continuously in the field.

What you’ll do

  • Develop evaluations for task success, answer quality, robustness, safety, cost, and the business outcome a system is intended to improve.
  • Build high-signal datasets, simulations, adversarial tests, and human-review protocols that capture what actually matters in a customer workflow.
  • Establish evaluation integrity by validating automated judges, monitoring drift, and making sure reported improvements represent real progress.
  • Investigate production traces and unexpected failures deeply enough to explain why they happened and which intervention is most likely to work.
  • Partner with forward deployed engineers to move research ideas into customer systems and measure their impact under real operating conditions.
  • Turn field results into reusable methods, internal standards, research infrastructure, and product capabilities across Verona.

What we’re looking for

  • A record of strong applied machine-learning research or engineering work, demonstrated through shipped systems, publications, open-source contributions, or equivalent projects.
  • Excellent experimental judgment: you can turn surprising system behavior into testable hypotheses and distinguish a meaningful result from a misleading metric.
  • Fluency in Python and the engineering ability to build durable research and evaluation infrastructure, not only one-off notebooks.
  • Strong foundations in statistics, machine learning, and the practical limitations of automated evaluation.
  • Clear technical communication and the ability to collaborate with researchers, production engineers, customer teams, and domain experts.
  • A bias toward research whose value can be observed in deployed systems and real user outcomes.

You might excel here if

  • Experience with LLM or agent evaluation, post-training, red teaming, synthetic data, reward modeling, or human-feedback systems.
  • Experience designing evaluations for open-ended workflows where quality is subjective, multidimensional, or difficult to observe directly.
  • Systems-level understanding of how models, retrieval, tools, prompts, data, and application code interact in production.
  • Experience working with messy domain data or subject-matter experts to turn tacit judgment into reliable measurement.

Benefits

  • Competitive compensation and equity
  • Generous health benefits
  • Unlimited PTO
  • Paid parental leave
  • Daily lunches and dinners
  • Transportation and relocation support
  • Retirement plans