CUDA Engineering Expert | $80-$100/hr Remote
Overview
We're looking for a GPU kernel optimization expert to join a project with a top AI lab. In this contract role, you'll dive deep into GPU kernels, using profiler-driven analysis to squeeze out maximum performance on modern hardware. If you live and breathe CUDA, C++17, and can reason about low-level parallelism, this is your chance to work with cutting-edge AI infrastructure.
What You'll Do5
- 1Profile and tune GPU kernels to boost performance, efficiency, and hardware utilization across different architectures.
- 2Leverage profiler metrics like L2 cache hit rate, throughput, and occupancy to pinpoint bottlenecks and guide optimization decisions.
- 3Review existing kernel code — even in unfamiliar algorithm domains — and identify performance issues without needing deep domain expertise.
- 4Write and modify GPU kernel code using C++17, Python, and shader languages (e.g., CUDA, HIP), and clearly document your optimization rationale.
- 5Analyze when specific profiler signals are (or aren't) useful, and communicate trade-offs effectively.
Requirements7
- 1Available for at least 20 hours per week on a contract basis.
- 2Strong command of C++17 — fluency with modern C++ features is a must.
- 3Working knowledge of Python and Git for collaborative development.
- 4Fluency in at least one GPU programming model: CUDA, HIP, HLSL, GLSL, or similar kernel programming frameworks.
- 5At least 1 year of professional or graduate-level research experience working with GPU architectures and optimization.
- 6Ability to use GPU profilers (e.g., NVIDIA Nsight Compute) to drive kernel improvements without requiring full algorithm context.
- 7Nice-to-have: experience with inline PTX assembly, tensor core optimization, NVIDIA Blackwell, or contributions to open-source GPU kernel libraries.
Who Should Apply
This role is for a hands-on performance engineer who enjoys the challenge of making GPU kernels run faster. You're comfortable reading assembly-level code, know your way around profiler dashboards, and can articulate why a kernel is memory-bound or compute-bound. You're also happy working independently on a remote contract basis, with a knack for documenting your work. Bonus points if you've shipped CUDA libraries or worked at GPU hardware companies like NVIDIA, AMD, or Qualcomm.
Salary Insight
The pay is $80–$100 per hour, based on experience and skill level. This is a contract position with flexible hours.
Application Tip
When applying, highlight a specific GPU kernel optimization project where you used profiler metrics to achieve a measurable speedup. Include before/after numbers (e.g., latency, throughput) to demonstrate your impact.
Similar open positions
Explore active roles that match your skills and interests.
Micro1
VerifiedComputational Engineering Expert | $40-$60/hr Remote
This role is your chance to shape the next generation of AI by feeding it real-world engineering expertise. You'll apply deep knowledge of computational simulation and systems engineering to train, evaluate, and improve how AI models tackle complex tasks like CFD, FEA, and robotics. It's a remote, project-based position where your feedback directly influences model performance.
Mercor
VerifiedPerformance Engineer (C++, Python, Rust) | $70-$110/hr Remote
We're looking for a Performance Engineer with deep expertise in low-level systems optimization to help train and evaluate cutting-edge Large Language Models. In this fully remote role, you'll use C++, Python, and Rust to design performance-engineering tasks, write solutions, and build evaluation rubrics that improve the quality of training data for frontier AI systems. This is a 40-hour-per-week W-2 contract (placed at a leading AI lab) with a competitive hourly rate.
Micro1
VerifiedComputational Physics Expert | $40-$60/hr Remote
This role sits at the intersection of computational physics and artificial intelligence. You’ll apply your deep scientific expertise to help train next-generation AI systems, ensuring they reason accurately and learn from high-quality real-world physics data. Your work will directly influence how models handle complex simulations, numerical modeling, and scientific analysis.
Mercor
VerifiedComputer Vision Expert | $80-$110/hr Remote
A leading AI lab is seeking experienced computer vision practitioners to join their GenAI team on a part-time, remote basis. In this role, you'll design real-world vision challenges, generate reference solutions, and evaluate frontier models to identify capability gaps. Your expertise will directly shape the next generation of AI systems.
Micro1
VerifiedSoftware Engineer | $50-$100/hr Remote
This isn't your typical software engineering role—you'll apply your coding expertise to shape how next-generation AI systems learn and reason. Your day-to-day work involves building robust backend and frontend systems that provide high-quality training data for AI models. You don't need prior AI experience; your strong engineering foundation is what matters most.
Mercor
VerifiedMachine Learning & NLP Expert | $80-$110/hr Remote
Join a top-tier AI lab's generative AI team as a part-time remote Machine Learning and NLP Expert. You'll craft complex, real-world tasks to test and improve frontier models, author reference solutions, and analyze model performance. This role requires hands-on Python proficiency and deep expertise in modern ML and NLP methods, offering flexible remote work at approximately 20 hours per week.