MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling)
- Pay
- $90 – $120 / Hour
- Open to
Canada
United Kingdom
United States
- Apply
We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.
- Skills
- mlops
- cuda
- triton
- llm serving
- vllm
- performance profiling
- pytorch
- jax
What you'll do
Models are getting good at writing application code, but ML systems work (why a kernel is memory-bound, what a profiler trace is really showing, why a KV cache is thrashing) is still thin in their training data. This role pays infrastructure engineers to fill that gap for a leading AI lab.
The work is training data creation and evaluation across four areas:
- GPU kernels: CUDA, Triton or Pallas
- Profiling and trace analysis: Kineto, torch.profiler, Nsight, XLA or JAX profiler
- Debugging: distributed and accelerator-bound workloads
- Inference serving: vLLM, SGLang, TensorRT-LLM, Ray Serve, paged attention, continuous batching
In each, you design hard, realistic tasks, write correct solutions, evaluate other tasks and solutions with written feedback, and build rubrics for kernel optimisation, profiler interpretation, distributed reasoning and serving trade-offs. You also coordinate with other subject experts to keep the data consistent.
Who qualifies
- 2+ years of hands-on professional work in ML systems, ML infrastructure, model serving, or GPU and accelerator performance. The ad is clear this is a systems role, not applied modelling or data science
- Practical experience in at least one of the four areas above; more than one is a strong plus
- Production JAX and/or PyTorch; framework-level depth (custom ops, FSDP, DDP, DeepSpeed, Megatron, compiler work) is a plus
- Familiarity with A100, H100, B200 or TPU and the throughput, latency and memory trade-offs
- Demonstrable career progression
- 40 hours a week on weekdays, with no conflicting engagements
- Must be located in Canada, the United Kingdom or the United States
What it pays
$90–120 per hour. For a 2-year experience bar that is a strong rate, reflecting how scarce kernel and serving expertise is.
Check the employment terms
The listing is labelled an hourly contract and its footer carries Mercor's standard independent-contractor terms, including weekly payment on Stripe or Wise and the H-1B and STEM OPT exclusion. But the body describes a 40-hour W-2 employment position with Cincinnatus LLC, placed at the lab as extended workforce, with no other engagements allowed. W-2 status would also be unusual for UK and Canadian residents. Ask which terms apply to you before accepting.
Worth knowing
Good:
- High rate for a 2-year bar
- Open to Canada and the UK as well as the US
- Deep, technical work that uses real systems knowledge
Less good:
- Full-time, and exclusive: no other engagements
- Contradictory contractor and W-2 terms in the ad
- The lab is not named
- H-1B and STEM OPT holders are excluded per the footer
About this listing
Posted on Mercor, read on 24 September 2026. Pay, location requirement, qualifications and both sets of terms are quoted from the ad. See Mercor.
More roles at Mercor
See all 304Similar roles at other platforms
- Computer Vision Specialistmicro1 · $50 – $90 / Hour · 11d ago
- Data Analyst (AI Evaluation)micro1 · $30 – $60 / Hour · 14d ago
- Senior Software Engineer: AI Evaluation & BenchmarksAlignerr · $80 – $100 / Hour · 1mo ago
- Real-world terminal and infrastructure scenariosInvisible Technologies · $35 / Hour · 15d ago
Guides about Mercor
See all 52- How much does Mercor pay? Rates from 304 live listingsPay breakdown
- The Mercor AI interview: what happens, what it checks, and what comes afterInterview prep
- Mercor vs Alignerr: which to apply to, and how they differPlatform comparison
- Mercor, micro1, Outlier, Alignerr and Handshake AI compared: which to apply to firstPlatform comparison
Browse similar roles
Not the right fit?
See every open role, or get new ones on Telegram or Discord as they are added.