Experienced administrators design benchmark tasks for AI agents: realistic calendar, travel, expense and records scenarios with mock emails and receipts, the expert solution path, and a rubric of 35+ criteria. Five years of admin work in a mid-to-large organisation. Paid per accepted task, $20–40/hour band.
Remote AI training and data labeling jobs
Filter jobs
Location
- Worldwide74 jobs
- United States30 jobs
- United Kingdom6 jobs
- Canada3 jobs
- Indiano roles alongside your other filters
- Mexico1 job
Language
- English104 jobs
- Germanno roles alongside your other filters
- Spanishno roles alongside your other filters
- Frenchno roles alongside your other filters
- Japaneseno roles alongside your other filters
- Portugueseno roles alongside your other filters
Field: Business & Finance
- Languages & Linguistics119 jobs
- Audio & Voice114 jobs
- Engineering109 jobs
- Business & Finance107 jobs, applied. Activate to remove
- Software & IT88 jobs
- Health & Medicine78 jobs
- Law, Policy & Security70 jobs
- General & Data Collection68 jobs
- Science & Math65 jobs
- AI Safety & Evaluation63 jobs
- Video, Image & Design44 jobs
- Data, AI & ML35 jobs
- Writing & Education14 jobs
- Other fields6 jobs
Newest
107 open roles matching these filters · page 4 of 5
- $20 – $40 / HourWorldwide
Design finance Excel tasks from your own work (three-statement models, LBO and DCF valuation, forecasting), write model solutions, and evaluate AI attempts for a leading tech company's GenAI team. US only, W-2 through Cincinnatus LLC, 40 hours a week to end of September then 20. $70–100/hour.
$70 – $100 / HourOpen to United StatesJoin Mercor's bench of management consultants for future AI evaluation projects: writing grading criteria for consulting deliverables and scoring AI and human work. No project is open yet. For consultants with 1+ year at MBB or equivalent. Remote, $100–150/hour.
$100 – $150 / HourWorldwideDesign Excel tasks from cross-industry experience (data models, dashboards, Power Query, VBA automation), write the solutions, and evaluate AI attempts for a tech company's GenAI team. Breadth across functions, teaching or a PhD help. US only, W-2 via Cincinnatus LLC, 40 then 20 hours a week. $70–100/hour.
$70 – $100 / HourOpen to United StatesPaid pilot for US pharma forecasters: build or critique drug launch curves and judge whether forecast assumptions (analogs, ramp, peak share, loss of exclusivity) hold up, while writing rubrics that evaluate AI analysis of drugs. 5+ years, 10–20 hours over 1–2 weeks. $130–210/hour.
$130 – $210 / HourOpen to United StatesWork inside a leading AI lab's research team as its insurance and actuarial specialist: QA model outputs, write instruction specs and golden solutions, and build insurance benchmarks. Needs an FSA, ASA, FCAS or ACAS, or an active adjuster, underwriting or broking licence, plus 4+ years. Full-time W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesDesign marketing tasks and solutions, score AI outputs against rubrics, and refine marketing-specific evaluation guidelines for an AI lab. For marketers with 8+ years at top-tier brands or agencies and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesBuild retail merchandising, category and operations tasks, write solutions, and grade AI outputs against rubrics for an AI lab. For merchants, category managers and retail operators with 8+ years at major retailers and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesPaid pilot for US market access and pricing professionals: set and pressure-test gross-to-net and formulary-tier assumptions by drug class, judge whether coverage and rebating assumptions match real payer behaviour, and write rubrics for AI drug analysis. 5+ years, 10–20 hours. $175–200/hour.
$175 – $200 / HourWorldwideFull-time finance specialist embedded with a leading AI lab: QA model outputs, write instruction specs and golden solutions, and build finance benchmarks across FP&A, IB, asset management, PE, risk or treasury. 5+ years at a recognised institution, VP-level progression. W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesCoordinate the finance raters on a leading AI lab's training-data program: build tracking and escalation workflows, monitor throughput and quality, triage rater questions and keep finance tasks consistent. 5–10 years in finance or finance operations with team coordination experience. US only, 35+ hours a week, $40–60/hour.
$40 – $60 / HourOpen to United StatesEvaluate AI model outputs on underwriting, claims and risk reasoning against rubrics, design hard insurance tasks with worked solutions, and refine scoring guidelines. Needs 8+ years at a top-tier insurer or broker and prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesEvaluate AI finance outputs against rubrics, design hard finance tasks with worked solutions, and refine scoring guidelines for a leading AI lab. Needs 8+ years at a top-tier bank, asset manager or Big Four firm plus prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $65–90/hour.
$65 – $90 / HourOpen to United StatesQA AI-agent runs inside the Financial Forecaster planning app: check the agent used the right scenario, account, coordinate and basis, catch plausible-but-wrong answers, harden tasks and sharpen grading. Needs weekly hands-on Financial Forecaster use and 5+ years in FP&A, reporting, technical accounting or lender reporting. $70–110/hour, remote.
$70 – $110 / HourWorldwideComplete self-contained fund-ops exercises from mock ledgers, statements and notices: reconciliations, NAV variance attribution, trade-break investigations, corporate actions and payment exceptions, each graded against a rubric. 3+ years in fund admin or investment ops. US only, about 15 hours a week, $75–110/hour.
$75 – $110 / HourOpen to United StatesFull-time IB and M&A specialist embedded with a leading AI lab: vet model outputs on deal work, write instruction specs and golden solutions, and build finance benchmarks. 5+ years at a recognised institution, VP-level progression, MBA or CFA preferred. W-2 via Cincinnatus, hybrid Bay Area, $100–150/hour.
$100 – $150 / HourHybridOpen to United StatesGrade AI-generated slides, spreadsheets and documents for real-world finance quality, catching factual, visual and presentation errors and writing structured feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwideBuild a real estate deal model in Excel from a synthetic document pack, with live formulas and sourced assumptions, then score two other contractors' models against fixed criteria. For associates at real-estate-focused mid-market PE funds with ~2 years of IB first. US only, about 20–25 hours over 1.5–2 weeks, $80–100/hour.
$80 – $100 / HourOpen to United StatesFull-time PE and VC specialist embedded with a leading AI lab: QA model outputs on investment work, write instruction specs and golden solutions, and design finance benchmarks. 5+ years at a recognised institution with Principal or VP-level ownership of decisions. W-2 via Cincinnatus, hybrid Bay Area, $110–150/hour.
$110 – $150 / HourHybridOpen to United StatesTurn everyday insurance judgment into AI training data: design realistic underwriting, claims, actuarial and compliance scenarios, review AI outputs and write feedback. Open to every insurance specialty, 3+ years, US-based. $75/hour.
$75 / HourOpen to United StatesPrepare, audit and evaluate financial statements, advise on cost and budget systems, and apply accounting, economics and tax knowledge to AI training datasets. A bachelor's in accounting or a related field and a CPA, auditor or tax specialist title are required. One opening, contractor, $40–95/hour.
$40 – $95 / HourWorldwideFunds lawyers redline fund documents in simulated negotiations, grade how an AI handles them, and write the scoring criteria. Two years in the funds group of a corporate law firm, an ABA JD and active US bar admission are the stated bar. 100 openings, task-based pay at $90–130/hour and roughly 3.5 hours per task.
$90 – $130 / HourWorldwideBrokers, appraisers and property managers with 5–10+ years explain feasibility studies, valuations, leases and market trends in writing, and build realistic property management scenarios, to train AI. At least 10 hours a week. $40–65/hour, 10 openings; US, Canada and UK preferred, Europe and LATAM considered.
$40 – $65 / HourWorldwideBuild messy, realistic data-entry tasks (CSVs, PDFs, spreadsheets with seeded errors), define the correct end state, and write 35+ point rubrics that grade AI agents on them. Remote contractor, $20–35/hour, paid per accepted task. Suits people from healthcare claims, finance back office or legal ops.
$20 – $35 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.