ML systems engineers write and evaluate training tasks for a frontier lab across GPU kernels, performance profiling, distributed debugging and LLM inference serving, plus the rubrics that grade them. 2+ years of hands-on ML infrastructure work. Canada, UK or US; 40 hours a week. $90–120/hour.
Remote AI training and data labeling jobs
Filter jobs
Location: United States
Language
Field
- Languages & Linguistics2 jobs
- Audio & Voice4 jobs
- Engineering9 jobs
- Business & Finance24 jobs
- Software & IT14 jobs
- Health & Medicine14 jobs
- Law, Policy & Security8 jobs
- General & Data Collection7 jobs
- Science & Math5 jobs
- AI Safety & Evaluation2 jobs
- Video, Image & Design3 jobs
- Data, AI & ML5 jobs
- Writing & Educationno roles alongside your other filters
- Other fieldsno roles alongside your other filters
Newest
81 open roles matching these filters · page 3 of 4
- $90 – $120 / HourOpen to Canada, United Kingdom and 1 more country
Help an AI lab evaluate a performance transfer model: define what a faithfully captured performance looks like, curate easy and hard benchmark examples, and shape the reviewer pool. For character animators with 4+ years and 2+ feature or AAA credits. US, sessions in LA or NYC, $60–90/hour.
$60 – $90 / HourHybridOpen to United StatesAudit applied machine-learning tasks used to train and evaluate a frontier AI lab's models: experiment design, model selection, evaluation methodology, leakage and metric gaming. For US practitioners with 3+ years of hands-on experimental ML in PyTorch, TensorFlow, scikit-learn or XGBoost. $70–90/hour.
$70 – $90 / HourOpen to United StatesDesign Excel tasks from cross-industry experience (data models, dashboards, Power Query, VBA automation), write the solutions, and evaluate AI attempts for a tech company's GenAI team. Breadth across functions, teaching or a PhD help. US only, W-2 via Cincinnatus LLC, 40 then 20 hours a week. $70–100/hour.
$70 – $100 / HourOpen to United StatesCertified medical coders (CCDS, CHC, CCS or CPC) review clinical documentation and AI-generated notes for coding integrity, validate code-to-documentation alignment and annotate against guidelines. US-based, C1 in any second language, 10 hours a week. $45–65/hour.
$45 – $65 / HourOpen to United StatesPaid pilot for US pharma forecasters: build or critique drug launch curves and judge whether forecast assumptions (analogs, ramp, peak share, loss of exclusivity) hold up, while writing rubrics that evaluate AI analysis of drugs. 5+ years, 10–20 hours over 1–2 weeks. $130–210/hour.
$130 – $210 / HourOpen to United StatesEvaluate Neuron Kernel Interface (NKI) development tasks for a frontier AI lab: CUDA-to-NKI migration fidelity, Trainium performance optimisation and GPU-versus-Trainium numerical correctness. Requires 2+ years writing NKI kernels for Trainium or Inferentia2. US-only, $70–90/hour.
$70 – $90 / HourOpen to United StatesWork inside a leading AI lab's research team as its insurance and actuarial specialist: QA model outputs, write instruction specs and golden solutions, and build insurance benchmarks. Needs an FSA, ASA, FCAS or ACAS, or an active adjuster, underwriting or broking licence, plus 4+ years. Full-time W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesDesign marketing tasks and solutions, score AI outputs against rubrics, and refine marketing-specific evaluation guidelines for an AI lab. For marketers with 8+ years at top-tier brands or agencies and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesNative Korean speakers in Canada, South Korea, the UK or the US map the layout of real Korean PDF pages and transcribe every region exactly in Hangul (hanja and handwriting included) to train document AI. Remote hourly contract at $33.58/hour; every task is reviewed by a second Korean expert.
$33.58 / HourOpen to Canada, South Korea and 2 more countriesBuild retail merchandising, category and operations tasks, write solutions, and grade AI outputs against rubrics for an AI lab. For merchants, category managers and retail operators with 8+ years at major retailers and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesSenior US full-stack engineers (6+ years, end-to-end system ownership) build production-grade software on a leading AI lab's pre-release models, integrate tool interfaces and evaluation harnesses, and document model failure modes for researchers. Full-time W-2 through Cincinnatus LLC, 40 hours a week, $90–110/hour.
$90 – $110 / HourOpen to United StatesEvaluate GPU and accelerator kernel development tasks for a frontier AI lab: numerical correctness, benchmarking fairness, task scoping and compile or runtime validity across CUDA, Triton, NKI and Pallas. For US engineers with 3+ years of kernel work in at least two of those frameworks. $70–90/hour.
$70 – $90 / HourOpen to United StatesTurn ambiguous AI program requirements into clear, contradiction-free rater guidelines and rubrics across finance, retail, insurance, legal and sports. For linguists, instructional designers and technical writers with 3+ years and GenAI/RLHF guideline experience. US, 35+ hours a week, $45–65/hour.
$45 – $65 / HourOpen to United StatesAudit Kubernetes tasks used to train and evaluate a frontier AI lab's models: cluster-operations scenarios, manifest and Helm correctness, and failure-mode troubleshooting (CrashLoopBackOff, OOMKilled, eviction). For US engineers with 3+ years of production Kubernetes and Go, Python or TypeScript. $70–90/hour.
$70 – $90 / HourOpen to United StatesDesign adversarial prompts, find jailbreaks and policy failures, and document vulnerabilities in frontier AI models across cyber, biosecurity, fraud, misinformation and political content. Hourly remote contract at $70–84/hour for residents of Europe, the UK and the US; 5+ years' relevant experience required.
$70 – $84 / HourOpen to Albania, Austria and 38 more countriesFull-time finance specialist embedded with a leading AI lab: QA model outputs, write instruction specs and golden solutions, and build finance benchmarks across FP&A, IB, asset management, PE, risk or treasury. 5+ years at a recognised institution, VP-level progression. W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesCoordinate the finance raters on a leading AI lab's training-data program: build tracking and escalation workflows, monitor throughput and quality, triage rater questions and keep finance tasks consistent. 5–10 years in finance or finance operations with team coordination experience. US only, 35+ hours a week, $40–60/hour.
$40 – $60 / HourOpen to United StatesEvaluate vulnerability-reproduction and remediation tasks for a frontier AI lab: faithful CVE reproductions in Docker labs, sound fixes, and two-part verification (functionality plus vulnerability tests). For US AppSec engineers, pentesters and vulnerability researchers with 3+ years. $70–90/hour.
$70 – $90 / HourOpen to United StatesAuthor AI evaluation tasks from real drawing sets, documents and site photos, with the correct RFI response, coordination comments or markup as the answer. For licensed architects, project architects and job captains with 3+ years. US only, $45–60/hour.
$45 – $60 / HourOpen to United StatesEvaluate AI model outputs on underwriting, claims and risk reasoning against rubrics, design hard insurance tasks with worked solutions, and refine scoring guidelines. Needs 8+ years at a top-tier insurer or broker and prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesEvaluate AI finance outputs against rubrics, design hard finance tasks with worked solutions, and refine scoring guidelines for a leading AI lab. Needs 8+ years at a top-tier bank, asset manager or Big Four firm plus prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $65–90/hour.
$65 – $90 / HourOpen to United StatesCertified pharmacy technicians answer medication questions and review AI responses for an AI lab building prior authorization workflows: dosing, interactions, indications, PA requirements. US only, 30–40 hours a week during the project. A flat $35/hour.
$35 / HourOpen to United StatesComplete self-contained fund-ops exercises from mock ledgers, statements and notices: reconciliations, NAV variance attribution, trade-break investigations, corporate actions and payment exceptions, each graded against a rubric. 3+ years in fund admin or investment ops. US only, about 15 hours a week, $75–110/hour.
$75 – $110 / HourOpen to United States
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.