QA AI-agent runs inside the Financial Forecaster planning app: check the agent used the right scenario, account, coordinate and basis, catch plausible-but-wrong answers, harden tasks and sharpen grading. Needs weekly hands-on Financial Forecaster use and 5+ years in FP&A, reporting, technical accounting or lender reporting. $70–110/hour, remote.
Remote AI training and data labeling jobs
Filter jobs
Location
Language
Field
- Languages & Linguistics119 jobs
- Audio & Voice114 jobs
- Engineering109 jobs
- Business & Finance107 jobs
- Software & IT88 jobs
- Health & Medicine78 jobs
- Law, Policy & Security70 jobs
- General & Data Collection68 jobs
- Science & Math65 jobs
- AI Safety & Evaluation63 jobs
- Video, Image & Design44 jobs
- Data, AI & ML35 jobs
- Writing & Education14 jobs
- Other fields6 jobs
Newest
716 open roles · page 24 of 30
- $70 – $110 / HourWorldwide
Write or verify hard ten-option multiple-choice questions for an AI benchmark across clinical medicine, imaging, pharmacovigilance, health economics and rehabilitation, with step-by-step solutions and references. MD, DO, PhD or doctoral candidate, 10+ hours a week, asynchronous. $94–119/hour.
$94 – $119 / HourWorldwideWrite or verify 10-option multiple-choice benchmark questions in applied maths (signal processing, actuarial science, optimization, climate modeling and more), with chain-of-thought solutions and references. For maths PhDs and doctoral candidates. Remote, 10+ hours a week, $61–77/hour.
$61 – $77 / HourWorldwideAuthor executable scientific-computing problems in ecology, biochemistry and genetics for Sci Code, a new AI benchmark: source a paper, dataset or repo, write the prompt and grading criteria, and keep it only if frontier models mostly fail. PhD plus Python or R, Git and Docker. 6 weeks, 20+ hours a week, $70/hour.
$70 / HourWorldwideSenior materials scientists, and electrical or mechanical engineers, author realistic tasks with a prompt, a data room and a grading method, run them against an AI model and tighten them until the model can no longer reason through cleanly. Daily onboarding and office hours. Remote hourly contract at $60–90/hour.
$60 – $90 / HourWorldwideRecord yourself from a trailing third-person camera while walking, moving through changing scenery and driving, to build video datasets for AI. Unlike micro1's 360° Video Recorder listing, you must already own both the camera and the mount. $20/hour, 100 openings.
$20 / HourWorldwideComplete self-contained fund-ops exercises from mock ledgers, statements and notices: reconciliations, NAV variance attribution, trade-break investigations, corporate actions and payment exceptions, each graded against a rubric. 3+ years in fund admin or investment ops. US only, about 15 hours a week, $75–110/hour.
$75 – $110 / HourOpen to United StatesFull-time IB and M&A specialist embedded with a leading AI lab: vet model outputs on deal work, write instruction specs and golden solutions, and build finance benchmarks. 5+ years at a recognised institution, VP-level progression, MBA or CFA preferred. W-2 via Cincinnatus, hybrid Bay Area, $100–150/hour.
$100 – $150 / HourHybridOpen to United StatesRemote hourly contract for developers on any stack who already use Visual Studio Code on their own Mac. $55–65/hour, paid weekly via Stripe or Wise. You need your own Mac and a display above 2.5 megapixels. Tasks are not described in the ad.
$55 – $65 / HourWorldwideGrade AI-generated slides, spreadsheets and documents for real-world finance quality, catching factual, visual and presentation errors and writing structured feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwideBuild a real estate deal model in Excel from a synthetic document pack, with live formulas and sourced assumptions, then score two other contractors' models against fixed criteria. For associates at real-estate-focused mid-market PE funds with ~2 years of IB first. US only, about 20–25 hours over 1.5–2 weeks, $80–100/hour.
$80 – $100 / HourOpen to United StatesFull-time PE and VC specialist embedded with a leading AI lab: QA model outputs on investment work, write instruction specs and golden solutions, and design finance benchmarks. 5+ years at a recognised institution with Principal or VP-level ownership of decisions. W-2 via Cincinnatus, hybrid Bay Area, $110–150/hour.
$110 – $150 / HourHybridOpen to United StatesTurn real lab and test-bench experience into AI benchmark tasks: build test logs, calibration records and failure reports, define the right diagnosis, and write 35+ point rubrics. Remote contractor, $30–70/hour paid per accepted task, 50 openings, 4+ years in test or production.
$30 – $70 / HourWorldwideAnnotate and score robotics video footage against a scoring guide: label actions, objects and events, rate clips for quality and relevance, and QA your own annotations. Mid-level annotation experience preferred, robotics footage experience heavily preferred. Flat $7/hour, 20 openings, remote contract.
$7 / HourWorldwideTag, segment and mark keyframes in video using Final Cut Pro on macOS to build AI training datasets. Professional or academic Final Cut experience counts; aimed at entry to mid-level editors. One opening, contractor, $15–80/hour.
$15 – $80 / HourWorldwideMechanical designers build and validate parametric FreeCAD models of parts and assemblies, applying GD&T, to create a CAD dataset for AI training. Optional Python scripting. Remote contractor role, 50 openings, $40–100/hour.
$40 – $100 / HourWorldwideTurn everyday insurance judgment into AI training data: design realistic underwriting, claims, actuarial and compliance scenarios, review AI outputs and write feedback. Open to every insurance specialty, 3+ years, US-based. $75/hour.
$75 / HourOpen to United StatesJudge Vietnamese audio clips, from human speakers and AI speech models, for tones, pronunciation and naturalness, and back each rating with a written English explanation. Native Vietnamese and B2+ English preferred. 100 openings, contractor, $30–65/hour.
$30 – $65 / HourWorldwideA full-time, salaried research role at micro1 designing benchmarks, rubrics, datasets and evaluation pipelines for frontier coding agents. Base salary $200,000–260,000 plus equity and benefits, remote, one opening. Three years in software engineering, ML or evaluation.
$200000 – $260000 / YearWorldwideListen to Bengali audio clips and rate how native and fluent the speaker or AI model sounds, then justify each rating in written English. Native Bengali plus B2 English; no AI experience needed. 100 openings, contractor, $30–65/hour.
$30 – $65 / HourWorldwideNative or near-native Telugu speakers transcribe audio and video, segment and label speech data, and proofread existing Telugu transcripts for an AI language data project, flagging dialect and cultural nuance. No AI experience needed. Ten openings, contractor, $10–24/hour.
$10 – $24 / HourWorldwideClinicians whose main practice is young abuse survivors review AI mental-health guidance on abuse cases, build case scenarios, and write trauma-informed best-practice content so models respond safely to vulnerable youth. MD or PhD (psychology or social work) preferred. $100–200/hour, 8 openings, contractor, remote.
$100 – $200 / HourWorldwideA part-time fellowship for federal civil litigators: draft and evaluate motions, judge where AI-written advocacy falls short of persuasive, and build the grading criteria that measure it. Requires at least three documented federal motion wins. 50 openings, $150–300/hour, remote contractor.
$150 – $300 / HourWorldwideReview AI-generated and human-written text for grammar, usage, meaning and instruction-following, then write short, objective feedback against detailed guidelines. Native U.S. English is required. Fifty openings, contractor, $40–50/hour, at least 20 hours a week on a four-week initial sprint.
$40 – $50 / HourWorldwide
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.