Skip to content
Labeling Jobs

Trainium (NKI) Kernel Expert

Pay
$70 – $90 / Hour
Open to
  • United States
Apply

We earn a commission if you sign up through the links on this page. It costs you nothing and does not affect which jobs we list. How this works.

Skills
  • nki
  • aws trainium
  • kernel optimization
  • cuda
  • performance profiling
  • numerical methods

What you'll do

This is the narrowest listing in Mercor's current group of $70–90 technical auditor roles. The tasks involve writing and porting kernels for AWS Trainium through the Neuron Kernel Interface (NKI), and a frontier lab uses them to train and evaluate its models. You judge whether each task, and its reference solution, is correct and appropriate for the hardware.

Three areas get the most scrutiny:

  • Migration fidelity: when a CUDA kernel is ported to NKI, does it compute the same thing, and does it use Trainium idioms rather than a line-by-line GPU transliteration?
  • Performance quality: is the optimisation actually right for NeuronCore, in tiling, memory placement across SBUF, PSUM and HBM, partition-dimension constraints and DMA orchestration?
  • Numerical correctness: are cross-platform tolerances sensible given differences in accumulation order, rounding and mixed-precision behaviour between GPU and Trainium?

Feedback is written against a rubric.

Who fits

Basic qualifications:

  • 2+ years developing or optimising NKI kernels for Trainium or Inferentia2
  • Solid NKI patterns: tile-based computation, SBUF/PSUM/HBM hierarchy management, partition constraints, DMA orchestration
  • Experience assessing CUDA-to-NKI migration quality
  • Trainium profiling: NeuronCore pipeline utilisation, tensor-engine throughput, memory-bandwidth bottlenecks
  • Experience defining or judging cross-platform numerical-correctness standards

Preferred: Neuron SDK or compiler internals work, contributions to NKI kernel libraries, prior CUDA or Triton kernel development, knowledge of NeuronCore-v2 architecture and supported data types (FP32, BF16, FP8, INT8), and benchmarking on Trn1 or Trn2 instances.

Realistically, the people who meet this bar have worked at AWS Annapurna Labs, on a Neuron-adopting ML team, or at a company that has trained at scale on Trainium. It is a very small pool, which is the main argument for applying if you are in it.

What it pays

$70–90 per hour, weekly via Stripe or Wise. H-1B and STEM OPT candidates cannot be supported. Hours and duration are not published.

The broader GPU Kernel Expert listing pays the same and accepts NKI as one of several frameworks.

Worth knowing

Good:

  • Only 2 years of experience required, lower than sibling listings, reflecting how new NKI is
  • Very little competition for a highly specialised skill set
  • Asynchronous, remote

Less good:

  • The same $70–90 as generalist auditor listings, which undervalues rare NKI expertise
  • US-only, with the visa exclusion
  • No hours or duration published
  • Independent contractor; projects can be extended, shortened or ended early

About this listing

Posted by Mercor as a remote hourly contract for US residents, confirmed open on 24 September 2026. Qualifications and pay are the ad's own. See Mercor.

More roles at Mercor

See all 304

Similar roles at other platforms

Guides about Mercor

See all 52

Browse similar roles

Not the right fit?

See every open role, or get new ones on Telegram or Discord as they are added.