
CUDA Engineering Expert | $60-$100/hr Remote
Overview
A remote contract role focused on optimizing GPU kernels with CUDA for a customer project in collaboration with a top AI lab. You’ll profile, tune, and refactor CUDA and C++ code to boost throughput on modern GPUs. Experience with GLSL and WebGPU helps you implement shader logic and graphics workflows within existing pipelines. Clear, actionable documentation and technical communication are essential as you contribute to design discussions and stay current on GPU programming advances.
What You'll Do7
- 1Analyze and profile GPU kernels using CUDA and profiling tools to extract higher throughput on current hardware.
- 2Identify kernel bottlenecks with stakeholders and propose targeted optimization strategies.
- 3Refactor CUDA and C++ code to improve maintainability and adaptability across GPU architectures.
- 4Implement shader logic and graphics workflows with GLSL and WebGPU within existing pipelines.
- 5Document findings, steps taken, and performance gains in clear reports and technical notes.
- 6Contribute to design discussions on new GPU-based approaches and performance metrics.
- 7Keep up to date with GPU programming trends and share relevant insights to improve project outcomes.
Requirements6
- 1Proven CUDA programming experience with a track record of tuning GPU kernels for performance.
- 2Advanced C++ development skills in high-performance computing contexts.
- 3Hands-on experience with GLSL and WebGPU for graphics and compute shader work.
- 4Familiarity with GPU profilers (e.g., Nsight, Visual Profiler) for guided optimization.
- 5Strong analytical skills to assess kernel performance across different hardware generations.
- 6Excellent written and verbal communication for precise documentation and reporting.
Who Should Apply
The ideal candidate brings deep CUDA expertise and a history of practical kernel optimization. They should be comfortable collaborating with researchers and engineers in a remote, cross-disciplinary setup. This role may be less suitable for someone without hands-on GPU programming or who lacks experience with profiling tools. Common fit signals include a proven optimization track record and strong documentation habits; red flags are limited CUDA experience or difficulty translating profiling results into concrete improvements.
Salary Insight
Pay range is $60 - $100 per hour; compensation is discussed at the offer stage.
Location
Required Skills
Application Tip
Highlight a specific kernel you optimized, including before/after throughput and the profiling tools you used to measure the improvement.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Mercor
VerifiedCUDA Engineering Expert
Mercor is putting together a team of GPU kernel specialists for a project backed by a top-tier AI lab. You'll dig into kernel code, use profiler data to spot performance bottlenecks, and make targeted performance improvements across modern GPU hardware. The role is a short-term contract, and you don't need to be an expert in every underlying algorithm to contribute. The work centers on CUDA and C++17 skills, plus a practical eye for profiler metrics like occupancy and cache throughput.

Micro1
VerifiedGPU Programming Software Engineer
This remote contract role puts your GPU expertise to work on large language model training. You will build and refine GPU kernels and shaders with CUDA, WebGPU, or GLSL, then profile them until they hit the performance targets. C++ handles the host-side logic that ties everything together. The project is output-based, so you earn per completed task that passes specs, with rates listed at $60 to $95 per hour. Expect a quick start: roles often fill in two days and first tasks begin within 24 to 48 hours after onboarding.

Mercor
VerifiedGPU Kernel Expert
This remote contract role puts you inside the evaluation loop for GPU/accelerator kernel tasks generated for a frontier AI lab. You will review assignments built on CUDA, Triton, NKI, and Pallas, checking whether they are numerically sound, correctly scoped, and safe to run. Your written, rubric-based feedback helps decide which tasks are used to train and evaluate the lab's models.

SME Careers
VerifiedC++ Engineer for AI Systems Remote Contract Role
As a remote C++ engineer, you will assess AI-generated code, system designs, and technical explanations to ensure quality. You will create clear, usable outputs by delivering reference implementations and detailed step-by-step reasoning for complex engineering problems. You will pinpoint issues in memory management, concurrency, and performance, validating solutions against prompts and documenting findings. This fully remote, hourly contractor role with SME Careers connects you to future AI data services projects and teams shaping foundation-model work.

Micro1
VerifiedComputational Engineering Expert
This role is your chance to shape the next generation of AI by feeding it real-world engineering expertise. You'll apply deep knowledge of computational simulation and systems engineering to train, evaluate, and improve how AI models tackle complex tasks like CFD, FEA, and robotics. It's a remote, project-based position where your feedback directly influences model performance.

Micro1
VerifiedSoftware Engineering Specialist Contract Remote Focused on Analysis and Writing
A remote contractor role that centers on technical analysis and documentation for a high-impact customer project. You’ll apply deep engineering knowledge to shape how models learn, reason, and perform by providing real-world input and clear rationale. No AI background required; your domain expertise is what matters, and you’ll produce thorough, accessible outputs for a broad audience. Software Engineering, System Design, and Technical Writing anchor the work as you help train next-generation AI systems.

