Skip to content
All roles
San Francisco, CA

AI Infrastructure Engineer

San Francisco, CAFull-timeSeniorOn-site

Zof AI is hiring for this role in San Francisco, CA. This is a full-time opportunity for candidates who want to contribute directly to the development of ambitious AI products in a high-performance environment.

Experience operating production AI or ML systems at scale is required.

Compensation

Competitive salary

Plus meaningful equity

About this role

Zof AI is seeking an AI Infrastructure Engineer to own the platform our AI systems run on. This is a consolidated platform role spanning what the market posts as AI Infrastructure, MLOps, LLMOps, and agent platform engineering: model serving and scaling, GPU and compute efficiency, and the provisioning, observability, and reliability infrastructure that keeps production agents stable. The ideal candidate has operated real AI workloads in production and builds infrastructure that lets a small team run systems well above its weight.

Responsibilities

  • Design and operate model serving, scaling, and compute infrastructure.
  • Own GPU and inference efficiency, capacity, and cost.
  • Build the platform for provisioning, running, and scaling agents.
  • Build observability for AI systems: tracing, monitoring, and alerting.
  • Keep production AI systems stable, debuggable, and recoverable.
  • Automate deployment, rollback, and environment management.
  • Set reliability standards and incident practices for AI workloads.
  • Partner with engineers to make the platform fast to build on.

Requirements

  • Experience operating production AI, ML, or high-scale backend systems.
  • Strong infrastructure and systems engineering foundation.
  • Experience with cloud platforms, containers, and orchestration.
  • Experience with observability and reliability tooling.
  • Judgment about cost, performance, and operational trade-offs.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership of systems in production.

Nice to have

  • Experience with GPU clusters, inference servers, or model gateways.
  • Experience running agent workloads or long-lived AI processes.
  • Experience with Kubernetes, Terraform, or similar tooling.
  • Experience in early-stage platform teams.
AI InfrastructureMLOpsLLMOpsAgent PlatformReliability
Apply

Apply for AI Infrastructure Engineer

This application is for the AI Infrastructure Engineer role in San Francisco, CA. Please complete every required field.

You are applying for

AI Infrastructure Engineer

San Francisco, CA

Basic information

City and country

Professional links

A link to Google Drive, Dropbox, LinkedIn, your website, or a PDF

Role-specific background

0/3000

Location fit

Written responses

0/3000
0/3000
0/3000
0/3000

For example, describe a project where you used AI tools to build faster, improve quality, or solve a difficult technical problem.

0/3000
0/3000

Final confirmation

By submitting, you agree that Zof AI may store and review your application for hiring purposes.

Related roles

Accra, GhanaFull-timeSeniorOn-site

AI Infrastructure Engineer

Zof AI is seeking an AI Infrastructure Engineer to own the platform our AI systems run on. This is a consolidated platform role spanning what the market posts as AI Infrastructure, MLOps, LLMOps, and agent platform engineering: model serving and scaling, GPU and compute efficiency, and the provisioning, observability, and reliability infrastructure that keeps production agents stable. The ideal candidate has operated real AI workloads in production and builds infrastructure that lets a small team run systems well above its weight.

GHS 8,000-16,000 / month

AI InfrastructureMLOpsLLMOpsAgent PlatformReliability
01Zof Console

One surface for posture, operations, and what needs attention next.

The authenticated home that engineering, QA, and SRE teams open every day: quality posture, in-flight runs, coverage by module, and what needs attention next.

OPERATIONAL KPIs

  • Runs
  • Coverage
  • Risk

Live across every environment you ship to.

WORK SPINE

  • Specs
  • Tests
  • Schedules

From specification to scheduled regression.

GUARDRAILS

  • RBAC
  • SSO
  • audit

Every action attributable to a named human.

LIVE/console
Zof AI home command center showing 12 runs at 94% pass, 3 open critical issues, 84% coverage, four module traceability bars, the specification pipeline, upcoming schedules, and recommended next actions with an active-runs sidebar.
Console home · Checkout Service · Staging · captured live from the product.
AI Infrastructure Engineer in San Francisco, CA | Careers at Zof AI