
GPU Kernel Expert | $70-$90/hr Remote
Overview
This remote contract role puts you inside the evaluation loop for GPU/accelerator kernel tasks generated for a frontier AI lab. You will review assignments built on CUDA, Triton, NKI, and Pallas, checking whether they are numerically sound, correctly scoped, and safe to run. Your written, rubric-based feedback helps decide which tasks are used to train and evaluate the lab's models.
What You'll Do8
- 1Assess kernel tasks for numerical correctness by reviewing tolerance settings and reference-implementation choices.
- 2Run performance checks using profiling tools such as Nsight Compute, ncu, or framework-native profilers to confirm benchmarks are fair.
- 3Verify that each task compiles and executes without hidden runtime issues across target accelerator environments.
- 4Identify common kernel failures, including driver mismatches, out-of-memory states, launch-configuration errors, shape and stride mismatches, and autotuning failures.
- 5Evaluate tasks across multiple categories, including generation from specification, translation between frameworks, hardware migration, debugging, performance optimization, and operator fusion.
- 6Write clear, rubric-based feedback that explains quality scores and points to specific improvements for each kernel task.
- 7Flag any experimental setup that could bias performance comparisons before the task enters a training or evaluation pipeline.
- 8Document findings in a structured written format that a frontier AI lab can use to refine task generation.
Requirements9
- 13+ years of hands-on experience developing, optimizing, or verifying GPU or accelerator kernels in at least two of the following: CUDA, Triton, NKI, or Pallas.
- 2Deep understanding of numerical-correctness standards, including absolute, relative, and ULP tolerances, plus how to select a proper reference implementation.
- 3Proven experience with performance profiling and benchmarking tools like Nsight, ncu, roofline analysis, or framework-native profilers.
- 4Working knowledge of common compilation and runtime problems, such as driver version mismatches, OOM errors, launch configuration failures, shape or stride mismatches, and autotuning inconsistencies.
- 5Practical exposure to at least three kernel task types: generation from specification, translation or lowering, migration across hardware targets, debugging, performance optimization, or operator fusion.
- 6Preferred: experience across both NVIDIA systems (CUDA or Triton) and custom accelerators such as NKI, Pallas, or TPU.
- 7Preferred: background in compiler engineering, MLIR, or intermediate-representation lowering.
- 8Preferred: familiarity with memory-hierarchy optimization, including shared-memory tiling, register pressure, bank conflicts, and coalescing patterns.
- 9Preferred: prior contributions to kernel libraries such as cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls.
Who Should Apply
The ideal candidate is an engineer who enjoys auditing and critiquing kernel implementations more than building them from scratch. If you prefer long hands-on coding sessions over written evaluation and feedback, this role will likely feel like a poor fit. Candidates often get rejected when they cannot show actual profiling results or when they struggle to explain ULP tolerances and reference-implementation reasoning. Another common gap is narrow experience, for example only writing CUDA kernels without touching translation, debugging, or operator fusion.
Salary Insight
The role pays $70 to $90 per hour, which is a strong contract rate for a kernel specialist with 3+ years of experience and multi-framework profici. The final rate depends on your breadth across GPU and custom-accelerator ecosystems and how well you can demonstrate evaluation skills.
Location
Required Skills
Application Tip
Include a short audit report you have written for a real kernel task, especially one that shows how you identified a benchmark fairness issue or a numerical tolerance flaw.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedCUDA Engineering Expert
Mercor is putting together a team of GPU kernel specialists for a project backed by a top-tier AI lab. You'll dig into kernel code, use profiler data to spot performance bottlenecks, and make targeted performance improvements across modern GPU hardware. The role is a short-term contract, and you don't need to be an expert in every underlying algorithm to contribute. The work centers on CUDA and C++17 skills, plus a practical eye for profiler metrics like occupancy and cache throughput.

Micro1
VerifiedCUDA Engineering Expert
A remote contract role focused on optimizing GPU kernels with CUDA for a customer project in collaboration with a top AI lab. You’ll profile, tune, and refactor CUDA and C++ code to boost throughput on modern GPUs. Experience with GLSL and WebGPU helps you implement shader logic and graphics workflows within existing pipelines. Clear, actionable documentation and technical communication are essential as you contribute to design discussions and stay current on GPU programming advances.

Micro1
VerifiedGPU Programming Software Engineer
This remote contract role puts your GPU expertise to work on large language model training. You will build and refine GPU kernels and shaders with CUDA, WebGPU, or GLSL, then profile them until they hit the performance targets. C++ handles the host-side logic that ties everything together. The project is output-based, so you earn per completed task that passes specs, with rates listed at $60 to $95 per hour. Expect a quick start: roles often fill in two days and first tasks begin within 24 to 48 hours after onboarding.

Mercor
VerifiedTrainium (NKI) Kernel Expert
You will evaluate the quality and correctness of NKI development tasks for a frontier AI lab's model training. Your reviews cover CUDA to NKI migration fidelity, Trainium-specific performance tuning quality, and cross-platform numerical-correctness standards. You will produce rubric-based written feedback that shapes how the lab builds and trains its models on AWS Trainium hardware.

Mercor
VerifiedExpert Interviewer
Mercor supports a frontier AI lab's high-priority technical-expert panel. The interviewer in this role vets shortlisted engineers across GPU kernel development, security & vulnerability research, and ML-compiler / accelerator domains. You run structured 20-minute interviews, score each candidate against a rubric, and provide an evidence-based read. The engagement is part of an AI model training-and-evaluation effort and operates 100% remote.

Mercor
VerifiedKubernetes Task Auditor
At Mercor, you will assess Kubernetes tasks that train a frontier AI lab's models. The work involves grading cluster-operation scenarios, checking manifest correctness, and judging whether failure-mode troubleshooting is realistic and complete. Each review ends with rubric-based written feedback that the AI team uses to improve model performance. This remote hourly role is open to candidates across the United States.

