Mercor
MercorVerified listing
Remote

Trainium (NKI) Kernel Expert | $70-$90/hr Remote

70–90/hr
Remote · Remote — United States
Posted August 26, 2026
hourly
3 openings
Upload your resume to see every role you match

Listing checked August 29, 2026 · pay as published by Mercor

Overview

You will evaluate the quality and correctness of NKI development tasks for a frontier AI lab's model training. Your reviews cover CUDA to NKI migration fidelity, Trainium-specific performance tuning quality, and cross-platform numerical-correctness standards. You will produce rubric-based written feedback that shapes how the lab builds and trains its models on AWS Trainium hardware.

What You'll Do7

  • 1Review NKI kernel development tasks for correctness, hardware fit, and performance quality.
  • 2Assess CUDA to NKI migration efforts for fidelity to the original kernel's behavior and performance.
  • 3Identify Trainium-specific bottlenecks such as NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth constraints.
  • 4Provide rubric-based written feedback that documents strengths, issues, and required changes.
  • 5Verify cross-platform numerical correctness, including accumulation order, rounding behavior, and mixed-precision semantics between GPU and Trainium.
  • 6Evaluate whether tasks respect NKI patterns like tile-based computation, SBUF/PSUM/HBM memory hierarchy, partition-dimension constraints, and DMA orchestration.
  • 7Collaborate with the AI lab to ensure kernel tasks align with Trainium hardware capabilities.

Requirements9

  • 12+ years of hands-on kernel development or tuning with NKI on AWS Trainium or Inferentia2.
  • 2Deep knowledge of NKI patterns: tile-based computation, SBUF/PSUM/HBM memory management, partition dimension constraints, and DMA orchestration.
  • 3Proven experience evaluating CUDA to NKI migration quality.
  • 4Familiarity with Trainium performance profiling, including NeuronCore pipeline utilization, tensor-engine throughput, and memory-bandwidth bottlenecks.
  • 5Experience defining or evaluating cross-platform numerical correctness standards, including mixed-precision semantics.
  • 6Experience with AWS Neuron SDK, Neuron Compiler internals, or NKI kernel libraries (preferred).
  • 7Prior CUDA or Triton kernel development (preferred).
  • 8Knowledge of Trainium hardware specifications such as NeuronCore-v2 architecture, on-chip SRAM, and supported data types FP32, BF16, FP8, INT8 (preferred).
  • 9Experience benchmarking ML training workloads on Trn1 or Trn2 instances (preferred).

Who Should Apply

The right expert has spent at least two years writing and tuning NKI kernels on AWS Trainium or Inferentia2. They can look at a CUDA kernel and judge whether its NKI translation preserves both accuracy and performance. This role is less suitable for CUDA or Triton developers who have not worked with NKI, because the review work centers on Trainium-specific patterns. Candidates often get rejected when they list NKI exposure but cannot describe SBUF/PSUM management or DMA orchestration in detail. Another common miss is lacking a method for checking cross-platform numerical correctness between GPU and Trainium.

Salary Insight

The contract pays $70 to $90 per hour. That rate matches senior kernel engineers who can review and evaluate NKI work, not just write kernels. Pay details beyond the hourly rate are not listed here.

Location

Typeremote
LocationRemote — United States
Eligible countriesUnited States
This is a remote position

Required Skills

nkicudatritonaws trainiuminferentia2aws neuron sdkneuron compilerneuroncoretrn1trn2fp32bf16fp8int8sramhbm

Application Tip

Share a specific NKI kernel tuning win from Trainium, including the profiling numbers before and after. Explain how you confirmed numerical correctness against a CUDA baseline to show you can handle the migration-fidelity review.

Share:

See NearSkill jobs more often in your search

How your application is processed

  1. 1Application received

    Your resume and details are logged the moment you apply.

  2. 2ATS + eligibility screening

    We check your profile against the role’s skills, seniority, and requirements.

  3. 3Employer sees qualified profiles only

    Only candidates who clear screening move forward.

See your fit score for every role

Similar open positions

Explore active roles that match your skills and interests.

Mercor

Mercor

10d agoRemotehourly

GPU Kernel Expert

This remote contract role puts you inside the evaluation loop for GPU/accelerator kernel tasks generated for a frontier AI lab. You will review assignments built on CUDA, Triton, NKI, and Pallas, checking whether they are numerically sound, correctly scoped, and safe to run. Your written, rubric-based feedback helps decide which tasks are used to train and evaluate the lab's models.

70–90/hr
· 3 openings
CudaTritonNki+10 more
Mercor

Mercor

1mo agoRemotetask-based

CUDA Engineering Expert

Mercor is putting together a team of GPU kernel specialists for a project backed by a top-tier AI lab. You'll dig into kernel code, use profiler data to spot performance bottlenecks, and make targeted performance improvements across modern GPU hardware. The role is a short-term contract, and you don't need to be an expert in every underlying algorithm to contribute. The work centers on CUDA and C++17 skills, plus a practical eye for profiler metrics like occupancy and cache throughput.

300 fixed
· 50 openings
C++PythonGit+10 more
Micro1

Micro1

25d agoRemotecontract
Hot

CUDA Engineering Expert

A remote contract role focused on optimizing GPU kernels with CUDA for a customer project in collaboration with a top AI lab. You’ll profile, tune, and refactor CUDA and C++ code to boost throughput on modern GPUs. Experience with GLSL and WebGPU helps you implement shader logic and graphics workflows within existing pipelines. Clear, actionable documentation and technical communication are essential as you contribute to design discussions and stay current on GPU programming advances.

60–100/hr
· 50 openings
CudaC++Glsl+1 more
SME Careers

SME Careers

5d agoRemotecontract

Kotlin Team Lead for Android and JVM QA and Training

A remote, hourly contract lead role focusing on Kotlin quality assurance and trainer performance across AI training projects. You’ll assess AI-generated Kotlin code, Android and JVM snippets, and coroutine workflows, ensuring alignment with project guidelines. Lead communications on Kotlin standards in Discord and guide contributors through onboarding materials and style guides. Expect to shape QA processes and provide precise feedback to keep outputs idiomatic, safe, and production-ready. Experience leading remote teams of trainers, engineers, and QAs is a plus.

Up to 65/hr
AIKotlinAndroid+23 more
SME Careers

SME Careers

5d agoRemotecontract

C++ Team Lead for AI Training Quality and QA

Lead the quality assurance effort for C++ AI training projects from a remote, contract role. Review AI-generated C++ code, assess compile-time and runtime behavior, and provide precise feedback to contributors. Maintain project guidelines, style guides, and rubrics, ensuring memory safety and performance. Collaborate with distributed teams of trainers, reviewers, and engineers to uphold production-ready quality. Strong C++ expertise and clear English communication are essential.

Up to 75/hr
C++CppMemory Management+30 more
Mercor

Mercor

11d agoRemotefull-time

Cloud / DevOps Engineer (Infra & IaC)

A leading AI lab building foundational Large Language Models needs a Cloud/DevOps Engineer to improve the quality of its training and inference infrastructure. The role centers on hands-on work with Kubernetes, AWS services, and Infrastructure-as-Code. You will design tasks, evaluate solutions, and write technical feedback that helps models reason about cluster failures, AWS integration, and CI/CD pipelines. The position is full-time (40 hours/week), remote within the United States, and structured as a W-2 contract through Cincinnatus LLC.

75–110/hr
· 10 openings
KubernetesAWSTerraform+6 more