Write and review Lean 4 proofs for a leading AI lab, formalize informal mathematics and judge whether a model's proof actually proves the right statement. A W-2 part-time employment position through Cincinnatus LLC, at least 20 hours a week and up to 40, paying $90–110/hour.
Remote AI training and data labeling jobs
Filter jobs
Location
- Worldwide6 jobs
- United States7 jobs
- United Kingdom1 job
- Canadano roles alongside your other filters
- Indiano roles alongside your other filters
- Mexicono roles alongside your other filters
Language
- English13 jobs
- Germanno roles alongside your other filters
- Spanishno roles alongside your other filters
- Frenchno roles alongside your other filters
- Japaneseno roles alongside your other filters
- Portugueseno roles alongside your other filters
Field: Software & IT
- Languages & Linguisticsno roles alongside your other filters
- Audio & Voiceno roles alongside your other filters
- Engineering29 jobs
- Business & Finance29 jobs
- Software & IT14 jobs, applied. Activate to remove
- Health & Medicine20 jobs
- Law, Policy & Security13 jobs
- General & Data Collection1 job
- Science & Math28 jobs
- AI Safety & Evaluation20 jobs
- Video, Image & Design3 jobs
- Data, AI & ML11 jobs
- Writing & Educationno roles alongside your other filters
- Other fieldsno roles alongside your other filters
Level: Senior
- Entryno roles alongside your other filters
- Junior1 job
- Medium17 jobs
- Senior14 jobs, applied. Activate to remove
Newest
14 open roles matching these filters
- $90 – $110 / HourWorldwide
Embedded firmware, FPGA/RTL, flight software and hardware test automation engineers review and write hard problems about software that runs on real hardware (timing, race conditions, interfaces, bring-up) for an AI research initiative. 5+ years, 3 recent hands-on. Remote contract, $100–120/hour.
$100 – $120 / HourWorldwideA short, well-paid sprint for very senior software engineers: help a leading foundation-model lab improve its models on hard SWE tasks. 10+ years at top US tech firms, about 20 hours a week for 2–3 weeks. $150–210/hour by geography and level; the 2-hour vetting exercise is paid $100.
$150 – $210 / HourWorldwideTwo to three paid one-hour video interviews with Mercor about how enterprise security work is done and judged, feeding a benchmark for AI cyber defense agents. $125–175/hour, US-based, no prep and nothing to label or submit.
$125 – $175 / HourOpen to United StatesA hybrid, Bay Area-based W-2 role embedded with a leading AI lab: senior software engineers vet model outputs, write instruction specs and golden solutions, and build engineering benchmarks. 4+ years, senior-level progression and a CS or engineering degree. 40 hours a week for an initial 6 months, $65–105/hour.
$65 – $105 / HourHybridOpen to United StatesA paid expert conversation for engineers who build and run LLM agents in production: a short AI screening interview (no coding), then, if selected, a 30-minute live call on agent reliability, evaluation and internal adoption, paid $100–500 depending on depth of experience.
$100 – $500 / TaskWorldwideAudit repository-level software engineering benchmark tasks for a frontier AI lab: reference patches, test harnesses, Docker isolation, and signs of answer leakage or reward hacking. For US engineers with 3+ years and real open-source contributor or maintainer history. $70–90/hour.
$70 – $90 / HourOpen to United StatesUS-based cloud and DevOps engineers design and grade AI training tasks on Kubernetes failure diagnosis, AWS service integration, Terraform or CDK design and CI/CD, and write the rubrics behind them. 4+ years at a top-tier organisation. Full-time W-2 through Cincinnatus LLC at a leading AI lab, $75–110/hour.
$75 – $110 / HourOpen to United StatesEvaluate Neuron Kernel Interface (NKI) development tasks for a frontier AI lab: CUDA-to-NKI migration fidelity, Trainium performance optimisation and GPU-versus-Trainium numerical correctness. Requires 2+ years writing NKI kernels for Trainium or Inferentia2. US-only, $70–90/hour.
$70 – $90 / HourOpen to United StatesSenior US full-stack engineers (6+ years, end-to-end system ownership) build production-grade software on a leading AI lab's pre-release models, integrate tool interfaces and evaluation harnesses, and document model failure modes for researchers. Full-time W-2 through Cincinnatus LLC, 40 hours a week, $90–110/hour.
$90 – $110 / HourOpen to United StatesEvaluate GPU and accelerator kernel development tasks for a frontier AI lab: numerical correctness, benchmarking fairness, task scoping and compile or runtime validity across CUDA, Triton, NKI and Pallas. For US engineers with 3+ years of kernel work in at least two of those frameworks. $70–90/hour.
$70 – $90 / HourOpen to United StatesAn expert-interview listing for engineers who have shipped production search, especially agentic search in the LLM era: a 25-minute conversational interview about relevance, evaluation and real trade-offs, with a possible paid 30-minute follow-up call at $200. No coding, no take-home. Listed at $80–150 per task.
$80 – $150 / TaskWorldwideGrade AI-generated slides, spreadsheets and documents for real-world software engineering quality, flagging factual, visual and presentation errors with written feedback. Needs 5+ years at a top firm in the US, UK, Canada, Australia or New Zealand. $100–150/hour.
$100 – $150 / HourWorldwideBuild and document the production pipelines you'd work on during a normal week, alongside experts from two other disciplines, so the outputs can be turned into the tasks and rubrics that train frontier models. $140–200/hour for 40 hours a week from 14 September to 10 October 2026. You must be based in the UK with the right to work there.
$140 – $200 / HourOpen to United Kingdom
Nothing that fits today?
New roles land most days. Get them on Telegram or Discord as they are added, or read how the platforms pay before you apply.