Predict when a named sell-side analyst will publish after a catalyst and what the note will say, reasoning only from a fixed evidence cutoff, then grade AI analyses of the same call. For former lead sell-side analysts or very senior associates in the exact sector, typically 8+ years. $150–250/hour, remote.
Remote AI training and data labeling jobs
Filter jobs
Location
- Worldwide30 jobs
- United States16 jobs
- United Kingdom2 jobs
- Canada1 job
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
- English45 jobs
- Germanno roles alongside your other filters
- Spanishno roles alongside your other filters
- Frenchno roles alongside your other filters
- Japaneseno roles alongside your other filters
- Portugueseno roles alongside your other filters
Field: Business & Finance
- Languages & Linguisticsno roles alongside your other filters
- Audio & Voiceno roles alongside your other filters
- Engineering55 jobs
- Business & Finance47 jobs, applied. Activate to remove
- Software & IT42 jobs
- Health & Medicine64 jobs
- Law, Policy & Security51 jobs
- General & Data Collection2 jobs
- Science & Math47 jobs
- AI Safety & Evaluation22 jobs
- Video, Image & Design6 jobs
- Data, AI & ML18 jobs
- Writing & Education4 jobs
- Other fields5 jobs
Newest
47 open roles matching these filters · page 2 of 2
- $150 – $250 / HourWorldwide
Read a specific non-G7 central bank's statements, minutes and speeches in the source language, score them dovish to hawkish from a fixed evidence cutoff, and grade AI macro analyses. For former central bank economists or senior local rates and FX strategists, typically 8+ years. $150–250/hour, remote.
$150 – $250 / HourWorldwideExperienced administrators design benchmark tasks for AI agents: realistic calendar, travel, expense and records scenarios with mock emails and receipts, the expert solution path, and a rubric of 35+ criteria. Five years of admin work in a mid-to-large organisation. Paid per accepted task, $20–40/hour band.
$20 – $40 / HourWorldwidePaid pilot for US pharma forecasters: build or critique drug launch curves and judge whether forecast assumptions (analogs, ramp, peak share, loss of exclusivity) hold up, while writing rubrics that evaluate AI analysis of drugs. 5+ years, 10–20 hours over 1–2 weeks. $130–210/hour.
$130 – $210 / HourOpen to United StatesWork inside a leading AI lab's research team as its insurance and actuarial specialist: QA model outputs, write instruction specs and golden solutions, and build insurance benchmarks. Needs an FSA, ASA, FCAS or ACAS, or an active adjuster, underwriting or broking licence, plus 4+ years. Full-time W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesDesign marketing tasks and solutions, score AI outputs against rubrics, and refine marketing-specific evaluation guidelines for an AI lab. For marketers with 8+ years at top-tier brands or agencies and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesBuild retail merchandising, category and operations tasks, write solutions, and grade AI outputs against rubrics for an AI lab. For merchants, category managers and retail operators with 8+ years at major retailers and prior LLM rubric experience. US, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesPaid pilot for US market access and pricing professionals: set and pressure-test gross-to-net and formulary-tier assumptions by drug class, judge whether coverage and rebating assumptions match real payer behaviour, and write rubrics for AI drug analysis. 5+ years, 10–20 hours. $175–200/hour.
$175 – $200 / HourWorldwideFull-time finance specialist embedded with a leading AI lab: QA model outputs, write instruction specs and golden solutions, and build finance benchmarks across FP&A, IB, asset management, PE, risk or treasury. 5+ years at a recognised institution, VP-level progression. W-2, hybrid Bay Area, $60–100/hour.
$60 – $100 / HourHybridOpen to United StatesCoordinate the finance raters on a leading AI lab's training-data program: build tracking and escalation workflows, monitor throughput and quality, triage rater questions and keep finance tasks consistent. 5–10 years in finance or finance operations with team coordination experience. US only, 35+ hours a week, $40–60/hour.
$40 – $60 / HourOpen to United StatesEvaluate AI model outputs on underwriting, claims and risk reasoning against rubrics, design hard insurance tasks with worked solutions, and refine scoring guidelines. Needs 8+ years at a top-tier insurer or broker and prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $60–80/hour.
$60 – $80 / HourOpen to United StatesEvaluate AI finance outputs against rubrics, design hard finance tasks with worked solutions, and refine scoring guidelines for a leading AI lab. Needs 8+ years at a top-tier bank, asset manager or Big Four firm plus prior hands-on LLM rubric evaluation. US only, 35+ hours a week, $65–90/hour.
$65 – $90 / HourOpen to United StatesQA AI-agent runs inside the Financial Forecaster planning app: check the agent used the right scenario, account, coordinate and basis, catch plausible-but-wrong answers, harden tasks and sharpen grading. Needs weekly hands-on Financial Forecaster use and 5+ years in FP&A, reporting, technical accounting or lender reporting. $70–110/hour, remote.
$70 – $110 / HourWorldwideFull-time IB and M&A specialist embedded with a leading AI lab: vet model outputs on deal work, write instruction specs and golden solutions, and build finance benchmarks. 5+ years at a recognised institution, VP-level progression, MBA or CFA preferred. W-2 via Cincinnatus, hybrid Bay Area, $100–150/hour.
$100 – $150 / HourHybridOpen to United StatesGrade AI-generated slides, spreadsheets and documents for real-world finance quality, catching factual, visual and presentation errors and writing structured feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwideFull-time PE and VC specialist embedded with a leading AI lab: QA model outputs on investment work, write instruction specs and golden solutions, and design finance benchmarks. 5+ years at a recognised institution with Principal or VP-level ownership of decisions. W-2 via Cincinnatus, hybrid Bay Area, $110–150/hour.
$110 – $150 / HourHybridOpen to United StatesBrokers, appraisers and property managers with 5–10+ years explain feasibility studies, valuations, leases and market trends in writing, and build realistic property management scenarios, to train AI. At least 10 hours a week. $40–65/hour, 10 openings; US, Canada and UK preferred, Europe and LATAM considered.
$40 – $65 / HourWorldwideExperienced IT leaders plan and run a small technology initiative with an external vendor, SLAs and stakeholder reporting, generating realistic management scenarios to train AI. Contractor, remote, 50 openings, $70–130/hour. Six years leading multi-quarter tech initiatives is the bar.
$70 – $130 / HourWorldwideRebuild client business slides in PowerPoint from images, improve their structure, storytelling and data visuals, and log every change with its rationale, producing training data for AI. Five years at an agency or brand team expected. 15 openings, $150–350/hour.
$150 – $350 / HourWorldwideBuild AI evaluation tasks in financial reporting, audit and technical accounting: scenarios, reference memos and workpapers, and rubrics that separate senior judgment from exam recall. US GAAP/PCAOB or IFRS/ISA track. 5+ years at a Big Four firm or as a corporate controller. $70–80/hour, remote.
$70 – $80 / HourWorldwideRefine CRM data models, arbitrate conflicting operational policies and document why you ruled the way you did, and design revenue workflows that hold up under vertical regulation. Four years with your hands actually in the system is the bar; the ad states outright that management-only exposure does not count. Remote contractor work.
$50 – $100 / HourWorldwideProduce the decks, models and operating-model work you'd build on a normal engagement, alongside experts from two other disciplines, so the outputs can be turned into the tasks and rubrics that train frontier models. $140–200/hour for 40 hours a week from 14 September to 10 October 2026. You must be based in the UK with the right to work there.
$140 – $200 / HourOpen to United KingdomDesign and review physician and patient questionnaires (screeners, question wording, response scales, branching logic, respondent burden) and say whether an instrument will actually produce usable data. A flat $140/hour, and therapeutic-area specialism is welcome but not required. Fully remote, on your own schedule.
$140 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.