Skip to content
Labeling Jobs

Real-world terminal and infrastructure scenarios

Pay
$35 / Hour
Open to
Worldwide
Apply

We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.

Skills
  • terminal
  • infrastructure
  • benchmark design
  • test automation
  • reference solutions
  • automated verifiers

What you'll do

You write tasks that AI agents will later be tested on. Each task is built around a realistic terminal or infrastructure scenario and has four parts, all named in the listing: the problem setup, the working environment the agent runs in, a reference solution that proves the task can be done, and an automated verifier that decides whether an attempt passed.

The listing asks for tasks that are both realistic and challenging, and the verifier has to be reliable. In practice that last part is where most of the effort goes. A verifier that accepts a wrong answer, or rejects a correct one done a different way, makes the task useless as a benchmark item, so expect to think hard about what "solved" means and how to check it from the outside.

The listing does not say which operating systems, tools or container formats the environments use, how many tasks you are expected to produce, or how tasks are reviewed.

Who fits

The listing publishes no formal requirements beyond the work itself. Reading the task description, the people who will do well here are those who:

  • Work in a shell every day and have fixed real infrastructure problems, not just followed tutorials
  • Can package a reproducible environment so a task behaves the same way every time it runs
  • Write tests that check outcomes rather than one exact sequence of commands
  • Can judge difficulty: hard enough to stretch an agent, not so obscure that it only tests trivia

Sysadmins, DevOps and SRE engineers, and backend developers who live in the terminal are the obvious fit. The listing does not state a years-of-experience bar or a country requirement.

What it pays

$35 an hour, a single fixed rate. That is roughly double the $17 most general Meridial listings pay, and well below the $60 of Invisible's Technical Task Auditor listing, which reviews tasks of this kind rather than writing them.

The listing's contract terms say payment is issued weekly via Stripe or Wise, and that Invisible cannot support H-1B or STEM OPT contractors.

Worth knowing

Good:

  • Real engineering work: designing environments and verifiers is closer to building test infrastructure than to rating chatbot answers
  • A clear, fixed hourly rate, above most of the marketplace
  • Fully remote, on your own schedule
  • Posted about a day before we read it, so it was fresh

Less good:

  • Two onboarding steps come first: a resume upload (about 5 minutes) and an unpaid qualification step Meridial estimates at about 30 minutes
  • No published requirements, so it is hard to judge your chances before applying
  • No minimum amount of work; projects can be shortened or ended early
  • The listing does not describe the review process or how many tasks are expected
  • Independent contractor terms: no benefits, your own tax to manage

About this listing

Read on the signed-in Meridial marketplace on 26 September 2026, posted about a day before we read it. The task description, rate and contract terms are the listing's own; the reading of who fits and the pay comparisons are ours. See Invisible Technologies.

More roles at Invisible Technologies

See all 26

Similar roles at other platforms

Guides about Invisible Technologies

See all 16

Browse similar roles

Not the right fit?

See every open role, or get new ones on Telegram or Discord as they are added.