Skip to content
All roles
San Francisco, CA

Evals Engineer

San Francisco, CAFull-timeMid to SeniorOn-site

Zof AI is hiring for this role in San Francisco, CA. This is a full-time opportunity for candidates who want to contribute directly to the development of ambitious AI products in a high-performance environment.

Must understand how to measure whether AI systems actually work, beyond demos.

Compensation

Competitive salary

Plus meaningful equity

About this role

Zof AI is seeking an Evals Engineer to build the tests that determine whether AI actually works. Verification is the heart of what Zof AI does, so this role sits close to the core of the product: designing eval suites, verification harnesses, and quality gates that confirm AI systems built the right thing, part QA discipline and part domain judgment. The ideal candidate is skeptical by default, rigorous about measurement, and motivated by turning "it seems to work" into evidence.

Responsibilities

  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.

Requirements

  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.

Nice to have

  • Experience building LLM evals, benchmarks, or test infrastructure.
  • QA, SDET, or test automation background.
  • Domain expertise in a vertical where correctness matters.
  • Experience with statistical evaluation methods.
EvalsVerificationAI QualityTestingQA
Apply

Apply for Evals Engineer

This application is for the Evals Engineer role in San Francisco, CA. Please complete every required field.

You are applying for

Evals Engineer

San Francisco, CA

Basic information

City and country

Professional links

A link to Google Drive, Dropbox, LinkedIn, your website, or a PDF

Role-specific background

0/3000

Location fit

Written responses

0/3000
0/3000
0/3000
0/3000

For example, describe a project where you used AI tools to build faster, improve quality, or solve a difficult technical problem.

0/3000
0/3000

Final confirmation

By submitting, you agree that Zof AI may store and review your application for hiring purposes.

Related roles

Accra, GhanaFull-timeMid to SeniorOn-site

Evals Engineer

Zof AI is seeking an Evals Engineer to build the tests that determine whether AI actually works. Verification is the heart of what Zof AI does, so this role sits close to the core of the product: designing eval suites, verification harnesses, and quality gates that confirm AI systems built the right thing, part QA discipline and part domain judgment. The ideal candidate is skeptical by default, rigorous about measurement, and motivated by turning "it seems to work" into evidence.

GHS 6,500-13,000 / month

EvalsVerificationAI QualityTestingQA
01Zof Console

Une surface pour la posture, les opérations et ce qui nécessite une attention particulière.

Le foyer authentifié que les équipes d'ingénierie, de QA et de SRE ouvrent chaque jour : posture de qualité, exécutions en vol, couverture par module et ce qui requiert de l'attention ensuite.

KPI OPÉRATIONNELS

  • Courses
  • Couverture
  • Risque

Vivez dans tous les environnements dans lesquels vous expédiez.

TRAVAIL DE LA Colonne Vertébrale

  • Spécifications
  • Tests
  • Horaires

De la spécification à la régression planifiée.

GARDE-CORPS

  • RBAC
  • SSO
  • audit

Chaque action attribuable à un humain nommé.

LIVE/console
Centre de commande domestique Zof AI affichant 12 exécutions à 94 % de réussite, 3 problèmes critiques ouverts, une couverture de 84 %, quatre barres de traçabilité des modules, le pipeline de spécifications, les calendriers à venir et les prochaines actions recommandées avec une barre latérale d'exécutions actives.
Vue d'accueil · Service de paiement · Mise en scène · capturé en direct à partir du produit.
Evals Engineer in San Francisco, CA | Careers at Zof AI