US-based cloud and DevOps engineers design and grade AI training tasks on Kubernetes failure diagnosis, AWS service integration, Terraform or CDK design and CI/CD, and write the rubrics behind them. 4+ years at a top-tier organisation. Full-time W-2 through Cincinnatus LLC at a leading AI lab, $75–110/hour.
Remote AI training and data labeling jobs
Filter jobs
Location: United States
Language: English
Field
- Languages & Linguistics1 job
- Audio & Voice4 jobs
- Engineering16 jobs
- Business & Finance30 jobs
- Software & IT16 jobs
- Health & Medicine16 jobs
- Law, Policy & Security19 jobs
- General & Data Collection10 jobs
- Science & Math5 jobs
- AI Safety & Evaluation2 jobs
- Video, Image & Design6 jobs
- Data, AI & ML7 jobs
- Writing & Education3 jobs
- Other fieldsno roles alongside your other filters
Newest
113 open roles matching these filters · page 4 of 5
- $75 – $110 / HourOpen to United States
Inpatient registered nurses with broad clinical exposure support ongoing clinical AI product work: annotation, clinical review, model evaluation, user-feedback investigation and guideline development, alongside engineers and clinicians. US only, 10 hours a week minimum, $55–65/hour.
$55 – $65 / HourOpen to United StatesDesign finance Excel tasks from your own work (three-statement models, LBO and DCF valuation, forecasting), write model solutions, and evaluate AI attempts for a leading tech company's GenAI team. US only, W-2 through Cincinnatus LLC, 40 hours a week to end of September then 20. $70–100/hour.
$70 – $100 / HourOpen to United StatesAudit end-to-end AI-assisted coding sessions (traces from tools like Cursor, Copilot or Claude Code) used to train and evaluate a frontier lab's models, judging correctness, workflow and reasoning with rubric-based feedback. 3+ years of software development plus hands-on agentic coding. US-only, $70–90/hour.
$70 – $90 / HourOpen to United StatesFull-time Bay Area hybrid role embedded with an AI lab: review model reasoning on materials problems, write golden solutions and specs, and build benchmarks. For materials PhDs (or master's with exceptional industrial depth) with 4+ years of R&D. W-2 via Cincinnatus, $70–110/hour.
$70 – $110 / HourHybridOpen to United StatesML systems engineers write and evaluate training tasks for a frontier lab across GPU kernels, performance profiling, distributed debugging and LLM inference serving, plus the rubrics that grade them. 2+ years of hands-on ML infrastructure work. Canada, UK or US; 40 hours a week. $90–120/hour.
$90 – $120 / HourOpen to Canada, United Kingdom and 1 more countryHelp an AI lab evaluate a performance transfer model: define what a faithfully captured performance looks like, curate easy and hard benchmark examples, and shape the reviewer pool. For character animators with 4+ years and 2+ feature or AAA credits. US, sessions in LA or NYC, $60–90/hour.
$60 – $90 / HourHybridOpen to United StatesAudit applied machine-learning tasks used to train and evaluate a frontier AI lab's models: experiment design, model selection, evaluation methodology, leakage and metric gaming. For US practitioners with 3+ years of hands-on experimental ML in PyTorch, TensorFlow, scikit-learn or XGBoost. $70–90/hour.
$70 – $90 / HourOpen to United StatesDesign Excel tasks from cross-industry experience (data models, dashboards, Power Query, VBA automation), write the solutions, and evaluate AI attempts for a tech company's GenAI team. Breadth across functions, teaching or a PhD help. US only, W-2 via Cincinnatus LLC, 40 then 20 hours a week. $70–100/hour.
$70 – $100 / HourOpen to United StatesCertified medical coders (CCDS, CHC, CCS or CPC) review clinical documentation and AI-generated notes for coding integrity, validate code-to-documentation alignment and annotate against guidelines. US-based, C1 in any second language, 10 hours a week. $45–65/hour.
$45 – $65 / HourOpen to United StatesPaid pilot for US pharma forecasters: build or critique drug launch curves and judge whether forecast assumptions (analogs, ramp, peak share, loss of exclusivity) hold up, while writing rubrics that evaluate AI analysis of drugs. 5+ years, 10–20 hours over 1–2 weeks. $130–210/hour.
$130 – $210 / HourOpen to United StatesEvaluate Neuron Kernel Interface (NKI) development tasks for a frontier AI lab: CUDA-to-NKI migration fidelity, Trainium performance optimisation and GPU-versus-Trainium numerical correctness. Requires 2+ years writing NKI kernels for Trainium or Inferentia2. US-only, $70–90/hour.
$70 – $90 / HourOpen to United StatesWork inside a leading AI lab's research team as its insurance and actuarial specialist: QA model outputs, write instruction specs and golden solutions, and build insurance benchmarks. Needs an FSA, ASA, FCAS or ACAS, or an active adjuster, underwriting or broking licence, plus 4+ years. Full-time W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesDesign marketing tasks and solutions, score AI outputs against rubrics, and refine marketing-specific evaluation guidelines for an AI lab. For marketers with 8+ years at top-tier brands or agencies and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesBuild retail merchandising, category and operations tasks, write solutions, and grade AI outputs against rubrics for an AI lab. For merchants, category managers and retail operators with 8+ years at major retailers and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesSenior US full-stack engineers (6+ years, end-to-end system ownership) build production-grade software on a leading AI lab's pre-release models, integrate tool interfaces and evaluation harnesses, and document model failure modes for researchers. Full-time W-2 through Cincinnatus LLC, 40 hours a week, $90–110/hour.
$90 – $110 / HourOpen to United StatesEvaluate GPU and accelerator kernel development tasks for a frontier AI lab: numerical correctness, benchmarking fairness, task scoping and compile or runtime validity across CUDA, Triton, NKI and Pallas. For US engineers with 3+ years of kernel work in at least two of those frameworks. $70–90/hour.
$70 – $90 / HourOpen to United StatesTurn ambiguous AI program requirements into clear, contradiction-free rater guidelines and rubrics across finance, retail, insurance, legal and sports. For linguists, instructional designers and technical writers with 3+ years and GenAI/RLHF guideline experience. US, 35+ hours a week, $45–65/hour.
$45 – $65 / HourOpen to United StatesAudit Kubernetes tasks used to train and evaluate a frontier AI lab's models: cluster-operations scenarios, manifest and Helm correctness, and failure-mode troubleshooting (CrashLoopBackOff, OOMKilled, eviction). For US engineers with 3+ years of production Kubernetes and Go, Python or TypeScript. $70–90/hour.
$70 – $90 / HourOpen to United StatesDesign adversarial prompts, find jailbreaks and policy failures, and document vulnerabilities in frontier AI models across cyber, biosecurity, fraud, misinformation and political content. Hourly remote contract at $70–84/hour for residents of Europe, the UK and the US; 5+ years' relevant experience required.
$70 – $84 / HourOpen to Albania, Austria and 38 more countriesFull-time finance specialist embedded with a leading AI lab: QA model outputs, write instruction specs and golden solutions, and build finance benchmarks across FP&A, IB, asset management, PE, risk or treasury. 5+ years at a recognised institution, VP-level progression. W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesCoordinate the finance raters on a leading AI lab's training-data program: build tracking and escalation workflows, monitor throughput and quality, triage rater questions and keep finance tasks consistent. 5–10 years in finance or finance operations with team coordination experience. US only, 35+ hours a week, $40–60/hour.
$40 – $60 / HourOpen to United StatesEvaluate vulnerability-reproduction and remediation tasks for a frontier AI lab: faithful CVE reproductions in Docker labs, sound fixes, and two-part verification (functionality plus vulnerability tests). For US AppSec engineers, pentesters and vulnerability researchers with 3+ years. $70–90/hour.
$70 – $90 / HourOpen to United StatesAuthor AI evaluation tasks from real drawing sets, documents and site photos, with the correct RFI response, coordination comments or markup as the answer. For licensed architects, project architects and job captains with 3+ years. US only, $45–60/hour.
$45 – $60 / HourOpen to United States
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.