
CUDA Engineering Expert | $500 One-Time Remote
Listing checked August 29, 2026 · pay as published by Mercor
Overview
Mercor is putting together a team of GPU kernel specialists for a project backed by a top-tier AI lab. You'll dig into kernel code, use profiler data to spot performance bottlenecks, and make targeted performance improvements across modern GPU hardware. The role is a short-term contract, and you don't need to be an expert in every underlying algorithm to contribute. The work centers on CUDA and C++17 skills, plus a practical eye for profiler metrics like occupancy and cache throughput.
What You'll Do6
- 1Inspect GPU kernels and refine them to improve execution speed, resource use, and overall hardware efficiency.
- 2Read profiler signals like L2 cache hit rate, L2 throughput, and occupancy to guide where and how to improve kernels.
- 3Find performance bottlenecks in kernel code without requiring deep familiarity with the algorithm being implemented.
- 4Write and edit code in C++17, Python, and GPU programming languages, and explain the reasoning behind changes.
- 5Use expertise in CUDA, HIP, shader languages, or similar to push kernel performance further.
- 6Keep clear records of performance improvement choices, including when specific profiler metrics helped or misled the process.
Requirements8
- 1Can commit at least 20 hours per week to the project.
- 2Comfortable with modern C++ up to C++17.
- 3Have working knowledge of Python and Git.
- 4Fluent in at least one GPU programming model (e.g., CUDA, HIP, Slang, HLSL, GLSL).
- 5At least one year of professional or graduate-level research experience working with GPUs.
- 6Know how to interpret GPU profiler metrics and turn them into kernel improvements.
- 7Can improve kernel performance even when you don't have deep familiarity with the underlying algorithm.
- 8Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX, tensor core tuning, NVIDIA Blackwell, or NSight Compute is a plus.
Who Should Apply
You're a GPU performance enthusiast who enjoys digging into profiler output and squeezing cycles out of kernels. You're comfortable working independently on a remote, contract basis, and you have a solid foundation in C++ and at least one GPU programming framework. You don't need to know every algorithm cold; you rely on profiling data and hardware knowledge to find wins.
Salary Insight
The listing references a one-time payment of $500 for this remote task-based project. No further compensation details were provided.
Pay and demand on NearSkill
Live dataMedian hourly pay, USD
$70/hour
Among 611 live similar roles that publish pay
- Live similar roles
- 835
- Listed in last 30 days
- 384
- Remote
- 97%
Hiring most right now: micro1 (455) · Mercor (122) · SME Careers (114)
Most requested skills
- python9%
- llm evaluation6%
- quality assurance5%
- Python5%
Figures from Data Engineering roles live on NearSkill when this page loaded. Only listings that publish USD pay are counted. A role can close before you apply, so check the listing itself.
Compare your resume against these live rolesLocation
Required Skills
Application Tip
When applying, highlight a specific kernel tuning project where you used profiler data (like NSight Compute) to make a measurable improvement. Include before-and-after metrics if you have them; that will show how you approach performance work.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Micro1
VerifiedCUDA Engineering Expert
A remote contract role focused on optimizing GPU kernels with CUDA for a customer project in collaboration with a top AI lab. You’ll profile, tune, and refactor CUDA and C++ code to boost throughput on modern GPUs. Experience with GLSL and WebGPU helps you implement shader logic and graphics workflows within existing pipelines. Clear, actionable documentation and technical communication are essential as you contribute to design discussions and stay current on GPU programming advances.

Mercor
VerifiedGPU Kernel Expert
This remote contract role puts you inside the evaluation loop for GPU/accelerator kernel tasks generated for a frontier AI lab. You will review assignments built on CUDA, Triton, NKI, and Pallas, checking whether they are numerically sound, correctly scoped, and safe to run. Your written, rubric-based feedback helps decide which tasks are used to train and evaluate the lab's models.

Micro1
VerifiedGPU Programming Software Engineer
This remote contract role puts your GPU expertise to work on large language model training. You will build and refine GPU kernels and shaders with CUDA, WebGPU, or GLSL, then profile them until they hit the performance targets. C++ handles the host-side logic that ties everything together. The project is output-based, so you earn per completed task that passes specs, with rates listed at $60 to $95 per hour. Expect a quick start: roles often fill in two days and first tasks begin within 24 to 48 hours after onboarding.

Mercor
VerifiedExpert Interviewer
Mercor supports a frontier AI lab's high-priority technical-expert panel. The interviewer in this role vets shortlisted engineers across GPU kernel development, security & vulnerability research, and ML-compiler / accelerator domains. You run structured 20-minute interviews, score each candidate against a rubric, and provide an evidence-based read. The engagement is part of an AI model training-and-evaluation effort and operates 100% remote.

Mercor
VerifiedComputational Structural & Mechanical Engineering Expert
We're hiring a computational structural and mechanical engineering expert to design challenging, graduate-level problems for evaluating advanced AI systems. In this role, you'll create original scientific workflows that force AI to genuinely use engineering software—running simulations, interpreting results, and planning experiments—rather than just pattern matching. You'll work with a stack of open-source tools like FEniCSx/DOLFINx, OpenFOAM, and MOOSE, and you'll iterate against cutting-edge AI models to fine-tune problem difficulty. This is a remote, hourly contract that rewards deep technical expertise and puzzle-design thinking.

Mercor
VerifiedEngineering & Software Domain Expert
Frontier AI models improve only when someone with real engineering experience checks their work. This role puts you inside a leading AI lab's GenAI team, where you will review engineering knowledge tasks, write instruction specs and golden solutions, and build benchmarks that measure model progress. You will work in the client's own tools, on-site in the Bay Area several days each week, and your output directly shapes how the model reasons about software and systems. Employment is W-2 through Cincinnatus LLC, with client-issued accounts and equipment.

