Skip to content
Labeling Jobs

GPU Kernel Expert

Pay
$70 – $90 / Hour
Open to
  • United States
Apply

We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.

Skills
  • cuda
  • triton
  • gpu kernels
  • performance profiling
  • jax pallas
  • numerical correctness

What you'll do

Kernel tasks are a growing part of how AI labs test coding models: write a fused operator from a spec, port a kernel between frameworks, fix a race, make something faster. You review those tasks for a frontier lab and judge whether each one is well built.

Four questions drive the review:

  • Numerical correctness: are the tolerances (absolute, relative or ULP) and the chosen reference implementation right for the operation and precision?
  • Benchmarking fairness: would the performance comparison actually reward a better kernel, or can it be won by measurement artefacts such as warm-up, caching or mismatched shapes?
  • Scoping: is the task specified tightly enough to have a clear solution?
  • Validity: will it compile and run, or does it trip over driver mismatches, OOM, bad launch configurations, shape or stride errors, or autotuning failures?

The ad lists the task types you may see: generation from a specification, translation or lowering across frameworks, migration between hardware targets, debugging, performance optimisation and operator fusion. You are expected to have worked on at least three of them.

Who fits

Basic qualifications:

  • 3+ years developing, optimising or verifying kernels in at least two of CUDA, Triton, NKI or Pallas (JAX)
  • Solid grasp of kernel numerical-correctness criteria
  • Profiling and benchmarking experience (Nsight, ncu, roofline analysis or framework profilers)
  • Familiarity with common compile and runtime failures
  • Experience with at least three of the task types above

Preferred: work across both NVIDIA and custom-accelerator ecosystems (NKI, Pallas, TPU), compiler or MLIR background, memory-hierarchy optimisation (shared-memory tiling, register pressure, bank conflicts, coalescing), and contributions to cuBLAS, cuDNN, Triton community kernels or JAX/XLA custom calls.

What it pays

$70–90 per hour, weekly via Stripe or Wise. H-1B and STEM OPT candidates cannot be supported. Hours and duration are not published.

Specialists whose depth is in Trainium specifically should also see the Trainium (NKI) Kernel Expert listing, which asks for 2+ years instead of 3+ but only on NKI.

Worth knowing

Good:

  • Broadest of Mercor's kernel listings: CUDA and Triton engineers qualify without accelerator experience, as long as they have a second framework
  • Benchmark-fairness review is interesting, high-judgment work
  • Asynchronous, remote

Less good:

  • The two-framework requirement excludes CUDA-only specialists
  • $70–90 is well below what senior kernel engineers command in industry
  • US-only, with the visa exclusion
  • Independent contractor; projects can be extended, shortened or ended early

About this listing

Posted by Mercor as a remote hourly contract for US residents, confirmed open on 24 September 2026. Qualifications and pay are the ad's own. See Mercor.

More roles at Mercor

See all 304

Similar roles at other platforms

Guides about Mercor

See all 52

Browse similar roles

Not the right fit?

See every open role, or get new ones on Telegram or Discord as they are added.