
Engineering & Software Domain Expert | $65-$105/hr Remote
Overview
Frontier AI models improve only when someone with real engineering experience checks their work. This role puts you inside a leading AI lab's GenAI team, where you will review engineering knowledge tasks, write instruction specs and golden solutions, and build benchmarks that measure model progress. You will work in the client's own tools, on-site in the Bay Area several days each week, and your output directly shapes how the model reasons about software and systems. Employment is W-2 through Cincinnatus LLC, with client-issued accounts and equipment.
What You'll Do5
- 1Evaluate engineering knowledge tasks and model responses for hidden flaws: missing behaviors, shallow reasoning, code that passes tests but fails code review, and unhandled edge cases.
- 2Write clear instruction specs and authoritative golden solutions for engineering problems, then define new tasks that mirror real practice.
- 3Design challenging engineering problems and evaluation sets, and co-build engineering-specific skills and tools with the research team.
- 4Align quality standards with client researchers and adjacent specialists, turning tacit engineering judgment into explicit criteria.
- 5Provide precise written feedback on model outputs and task quality.
Requirements9
- 1Bachelor's degree or higher in computer science, software engineering, or a related engineering field. MS or PhD strongly preferred.
- 24+ years building and shipping production software or engineered systems at a reputable company. Internships and coursework do not count.
- 3Deep specialization in at least one area such as distributed systems, security/cryptography, embedded/firmware, ML systems, developer tooling, platform engineering, or test engineering.
- 4Experience at a senior level: Senior, Staff, Principal, Architect, or Engineering Lead, with meaningful ownership of systems and design decisions.
- 5Proven engineering record through shipped systems, patents, open source contributions, or published technical work.
- 6Hands-on use of large language models in professional work, with the ability to distinguish solid reasoning from plausible-sounding errors.
- 7Ability to commit 40 hours per week for an initial 6-month engagement.
- 8Based in the Bay Area, or willing to relocate there at your own cost, and able to work on-site multiple days per week.
- 9Excellent written communication and the ability to give precise, well-structured feedback.
Who Should Apply
This role fits an experienced engineer who has shipped real systems, can defend design decisions, and already uses LLMs in their daily work. Candidates who are comfortable reviewing code for subtle correctness issues and writing precise evaluation criteria will thrive. The role is less suitable for generalists or engineers who have only done coursework and internships, since it demands 4+ years of production experience and a clear specialization. Applications often fail when the candidate cannot show senior-level ownership (titles alone are not enough) or when they are not prepared to relocate to the Bay Area at their own expense. A weak answer to 'how do you tell a well-reasoned model output from a plausible-wrong one' also disqualifies many applicants.
Salary Insight
The pay range is $65-$105 per hour. At the top end, this matches rates for senior engineers with deep niche expertise; at the lower end, it still sits above typical contract engineering rates. You will be W-2, so benefits and payroll are handled by Cincinnatus.
Location
Required Skills
Application Tip
In your application, include one concrete example of an engineering decision you owned that shipped to production, and describe how you used an LLM to evaluate or improve technical work. Show your domain depth with specific systems or tools you have mastered.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Micro1
VerifiedAI Software Engineering Domain Expert
Remote, part-time contract work with micro1 where you apply deep software engineering knowledge to help train next-generation AI systems. You’ll review and polish AI-generated technical content, refine prompts, and judge model outputs using rubrics to ensure accuracy and quality. The role centers on drafting and editing technical docs, architecture plans, RFCs, and design specs that feed AI training, plus independent research and data annotation. Strong writing and a solid engineering background are essential, and you’ll collaborate asynchronously with project leads to meet deliverables.

Mercor
VerifiedSoftware Engineer, Full Stack (Python, Java, Rust, C#, C++)
This role puts you inside a leading AI lab's generative AI team, where you'll build production-grade full-stack applications that push frontier large language models to their limits. You'll integrate unreleased model APIs into working software, uncover edge-case failures, and iterate quickly in a fast-paced research environment. The position is a full-time (40 hours/week) remote W-2 engagement through Cincinnatus LLC, with pay at $50–65 per hour.

Micro1
VerifiedComputational Engineering Expert
This role is your chance to shape the next generation of AI by feeding it real-world engineering expertise. You'll apply deep knowledge of computational simulation and systems engineering to train, evaluate, and improve how AI models tackle complex tasks like CFD, FEA, and robotics. It's a remote, project-based position where your feedback directly influences model performance.

Mercor
VerifiedFinance Domain Expert — AI Training & Evaluation
A leading AI lab's GenAI team needs a senior finance practitioner to judge and improve how frontier models handle real financial work. The role involves reviewing finance tasks and model outputs, writing instruction specs and golden solutions, and building benchmarks that measure model progress. This is a full-time W-2 position through Cincinnatus LLC, placed inside the client's own tools and workflows. The work is hybrid in the Bay Area, with on-site collaboration required several days each week.

Mercor
VerifiedEngineering Simulation Specialist
Collaborate with a leading AI research lab on an evaluation that tests whether frontier models can reason from first principles in engineering. You craft simulations and pass/fail spec sheets that the model must satisfy, spanning domains like control systems, analog circuits, and mechanical design. You get a direct view into the model's internal reasoning on your own tasks. The role runs through a browser-based studio and GitHub, with daily onboarding calls to get you up to speed.

Mercor
VerifiedSenior Software Engineer, Full Stack (Python, Java, Rust, C#, C++)
This role places senior full-stack engineers inside a top-tier AI lab's generative AI team, where you'll build real applications and internal tools on top of frontier large language models before they're publicly released. You'll work directly with the lab's engineering manager, moving quickly and shipping working code from day one. It's a fully remote, 40-hour-per-week W-2 contract through Cincinnatus LLC, with a competitive hourly rate.

