A hybrid, Bay Area-based W-2 role embedded with a leading AI lab: senior software engineers vet model outputs, write instruction specs and golden solutions, and build engineering benchmarks. 4+ years, senior-level progression and a CS or engineering degree. 40 hours a week for an initial 6 months, $65–105/hour.
Remote AI training and data labeling jobs
Filter jobs
Location: United States
- Worldwide238 jobs
- United States63 jobs, applied. Activate to remove
- United Kingdom9 jobs
- Canada5 jobs
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
Field
- Languages & Linguisticsno roles alongside your other filters
- Audio & Voiceno roles alongside your other filters
- Engineering6 jobs
- Business & Finance16 jobs
- Software & IT9 jobs
- Health & Medicine10 jobs
- Law, Policy & Security16 jobs
- General & Data Collection1 job
- Science & Math5 jobs
- AI Safety & Evaluation2 jobs
- Video, Image & Design3 jobs
- Data, AI & ML4 jobs
- Writing & Education2 jobs
- Other fieldsno roles alongside your other filters
Newest
63 open roles matching these filters · page 2 of 3
- $65 – $105 / HourHybridOpen to United States
Produce real design work in Figma (layouts, wireframes, prototypes, design systems) and critique design iterations in writing, as reference material for training AI on visual and UX quality. US-based designers with 5+ years and a strong portfolio. $30–80/hour, ten openings, contractor.
$30 – $80 / HourOpen to United StatesReview and annotate financial statements, critique audit procedures, build and check financial models, and assess tax compliance and legal-financial documents for AI training data. US-based contributors only; a CPA, auditor or tax specialist title and an accounting degree are required. Thirty openings, contractor, $60–120/hour.
$60 – $120 / HourOpen to United StatesFull-time Bay Area hybrid role with an AI lab's research team: review model reasoning on life sciences tasks, write golden solutions and instruction specs, and design benchmarks. For life sciences PhDs with 4+ years of substantive research experience. W-2 via Cincinnatus, $65–105/hour.
$65 – $105 / HourHybridOpen to United StatesUS-based systems analysts bring requirements gathering, BPMN or UML process modelling and SQL troubleshooting to the design and validation of AI training scenarios. Contractor, remote (US only), 10 openings, $60–120/hour. Five years as a systems analyst or technical consultant.
$60 – $120 / HourOpen to United StatesLitigate a simulated commercial contract dispute turn by turn: draft motions, briefs, demand letters and discovery, answer opposing counsel's moves, and write rubric items for what a strong response looks like. Three years of post-JD commercial contract litigation and a US licence (active or lapsed) are the profile. US-based, one opening, $100/hour.
$100 / HourOpen to United StatesFormer federal and military staff write and grade WARs, SITREPs, after-action summaries and executive briefings, supply golden reference answers, and help build the rubrics that train AI on federal report writing. Three years of DoD reporting preferred. US only, $40–80/hour, 20 openings.
$40 – $80 / HourOpen to United StatesAuthor AI evaluation tasks from real fire and life safety review work: egress plan checks, sprinkler and alarm review, hazmat control areas, firestop photo verification. For US fire marshals, fire protection engineers and NICET III+ designer-reviewers with 3+ years in the seat. Remote hourly contract at $45–60/hour.
$45 – $60 / HourOpen to United StatesAudit repository-level software engineering benchmark tasks for a frontier AI lab: reference patches, test harnesses, Docker isolation, and signs of answer leakage or reward hacking. For US engineers with 3+ years and real open-source contributor or maintainer history. $70–90/hour.
$70 – $90 / HourOpen to United StatesA full-time W-2 placement at a leading AI lab through Cincinnatus LLC: US-based mechanical engineers vet model outputs, write instruction specs and reference solutions, and build benchmarks. 5+ years in industry and a mechanical engineering degree required. 40 hours a week for an initial 2–3 months, $60–90/hour.
$60 – $90 / HourOpen to United StatesJudge AI-enhanced and upscaled video and stills at pixel level for a leading AI lab's GenAI team: artifacts, noise, aliasing, banding, grain, sharpening. For VFX and rendering supervisors, colorists, DPs, lighting artists and high-end photographers with 5+ years. US, 20 hours a week, $60–90/hour.
$60 – $90 / HourOpen to United StatesAct as ground truth for an AI lab teaching models real enterprise sales work: audit workflows, build golden reference trajectories in a mock sales stack, and refine rubrics. For sellers with around 10 years in enterprise sales. US only, $60–90/hour, placed via Cincinnatus.
$60 – $90 / HourOpen to United StatesPractising US primary care, family medicine or general internal medicine physicians review outpatient notes and judge AI-written documentation against what a PCP would actually chart. Needs C1 or better in one of 25 listed languages. 2+ years post-residency, 10 hours a week, $170–190/hour.
$170 – $190 / HourOpen to United StatesExperimental scientists create and review training data for a frontier lab's materials science models: inorganic synthesis, superconductors, semiconductors and advanced packaging, characterization (XRD, SEM, TEM) and fabrication. PhD, MS or equivalent hands-on experience. US-based, 10–40 hours a week, $84/hour.
$84 / HourOpen to United StatesFull-time W-2 role through Cincinnatus LLC, embedded with a leading AI lab: senior drug development scientists write golden solutions, instruction specs and benchmarks for pharma R&D reasoning. Hybrid in the Bay Area, 40 hours a week for an initial 6 months. PhD, PharmD or MD with 4+ years in industry R&D. $75–115/hour.
$75 – $115 / HourHybridOpen to United StatesA full-time W-2 role (via Cincinnatus LLC) embedded with a leading AI lab in the Bay Area: review legal model outputs, write instruction specs and golden solutions, and build legal benchmarks. Hybrid, on-site several days a week, 6-month initial term. $60–100/hour; JD, 5+ years' practice, US bar.
$60 – $100 / HourHybridOpen to United StatesFull-time, on-site-hybrid role in the Bay Area: review AI output on business and sales operations tasks, write instruction specs and golden solutions, and build benchmarks with an AI lab's research team. For senior ops leaders with 4+ years. W-2 via Cincinnatus, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesRedline MSAs, NDAs and DPAs against a company playbook, choose the position and the fallback, and write the reasoning behind every markup so an AI can learn it. US-based, with a JD, US bar admission and four years of transactional practice preferred. 50 openings at $90–110/hour.
$90 – $110 / HourOpen to United StatesUS-based cloud and DevOps engineers design and grade AI training tasks on Kubernetes failure diagnosis, AWS service integration, Terraform or CDK design and CI/CD, and write the rubrics behind them. 4+ years at a top-tier organisation. Full-time W-2 through Cincinnatus LLC at a leading AI lab, $75–110/hour.
$75 – $110 / HourOpen to United StatesFull-time Bay Area hybrid role embedded with an AI lab: review model reasoning on materials problems, write golden solutions and specs, and build benchmarks. For materials PhDs (or master's with exceptional industrial depth) with 4+ years of R&D. W-2 via Cincinnatus, $70–110/hour.
$70 – $110 / HourHybridOpen to United StatesHelp an AI lab evaluate a performance transfer model: define what a faithfully captured performance looks like, curate easy and hard benchmark examples, and shape the reviewer pool. For character animators with 4+ years and 2+ feature or AAA credits. US, sessions in LA or NYC, $60–90/hour.
$60 – $90 / HourHybridOpen to United StatesPaid pilot for US pharma forecasters: build or critique drug launch curves and judge whether forecast assumptions (analogs, ramp, peak share, loss of exclusivity) hold up, while writing rubrics that evaluate AI analysis of drugs. 5+ years, 10–20 hours over 1–2 weeks. $130–210/hour.
$130 – $210 / HourOpen to United StatesEvaluate Neuron Kernel Interface (NKI) development tasks for a frontier AI lab: CUDA-to-NKI migration fidelity, Trainium performance optimisation and GPU-versus-Trainium numerical correctness. Requires 2+ years writing NKI kernels for Trainium or Inferentia2. US-only, $70–90/hour.
$70 – $90 / HourOpen to United StatesWork inside a leading AI lab's research team as its insurance and actuarial specialist: QA model outputs, write instruction specs and golden solutions, and build insurance benchmarks. Needs an FSA, ASA, FCAS or ACAS, or an active adjuster, underwriting or broking licence, plus 4+ years. Full-time W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United States
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.