
Director of Infrastructure Engineering for Multi‑Cloud Platform
Overview
The role leads the architecture and evolution of a multi-cloud platform spanning AWS and GCP to support production AI systems. You will own the cloud strategy, developer experience, and reliability practices, ensuring security, observability, and resilience as the team grows. This is a hands-on leadership position that blends long‑term infrastructure planning with incident resolution and mentoring engineers. You’ll drive automation, platform tooling, and security-by-default across the organization, balancing speed with risk management.
What You'll Do8
- 1Own the multi-cloud infrastructure architecture for AWS and GCP, aligning it with scalability, reliability, security, and cost goals.
- 2Grow and coach a high‑performing Infrastructure/Platform Engineering team while fostering ownership and continuous improvement.
- 3Implement infrastructure as code with Terraform or equivalents to ensure reproducible, versioned, auditable environments.
- 4Build and evolve CI/CD pipelines to enable rapid, secure software delivery.
- 5Develop comprehensive observability using metrics, logs, traces, and alerts to detect and resolve production issues.
- 6Define reliability practices including incident response, on‑call handling, SLOs, error budgets, disaster recovery, and blameless postmortems.
- 7Collaborate with Security and Engineering leadership to embed security by default and maintain compliance with ISO 27001, SOC 2, and CMMC where applicable.
- 8Improve developer experience by reducing operational friction through automation and platform tooling.
Requirements8
- 18+ years building and operating production infrastructure, platform engineering, DevOps, or SRE systems.
- 23+ years leading engineering teams in high-growth settings.
- 3Deep experience with AWS, GCP, or multi‑cloud production environments.
- 4Strong command of Terraform, Kubernetes, containers, and modern CI/CD platforms.
- 5Proven track record delivering highly available, observable, and resilient production systems.
- 6Solid understanding of infrastructure security, compliance, and risk management.
- 7Experience scaling engineering and infrastructure in fast-moving startup environments.
- 8Excellent communication skills and ability to influence technical strategy with engineering and executive stakeholders.
Who Should Apply
The ideal candidate brings 8+ years of hands-on infrastructure leadership and has guided teams through rapid growth. You should be comfortable shaping cloud strategy across AWS and GCP, building scalable platforms, and championing security and reliability. This role is less suitable for someone who lacks multi‑cloud experience, cannot lead a team, or cannot align security/compliance with engineering goals. Common fit hurdles include limited experience with Terraform or Kubernetes, or difficulty communicating strategy to both engineers and executives.
Salary Insight
The posting lists a base pay range of 230,000–260,000 annually. Equity, bonuses, and remote-friendly benefits are available, with pay discussed during offers.
Location
Required Skills
Application Tip
Highlight Terraform, Kubernetes, and multi-cloud experience with concrete outcomes, such as reduced deployment time by a measurable percentage and improved SLO adherence to demonstrate reliability impact.
See NearSkill jobs more often in your search
How your application is processed
1Application received
Your resume and details are logged the moment you apply.
2ATS + eligibility screening
We check your profile against the role’s skills, seniority, and requirements.
3Employer sees qualified profiles only
Only candidates who clear screening move forward.
Similar open positions
Explore active roles that match your skills and interests.

Micro1
VerifiedTechnical Architect for Production Platform Engineering
A contractor role focused on platform engineering for next-generation AI systems. You’ll craft real-world, production-grade cloud scenarios to test how models learn, reason, and operate. Your guidance shapes how reinforcement learning environments evaluate design, deployment, security, and recovery across distributed infrastructure. Prior AI experience isn’t required; deep hands-on production expertise and domain knowledge take precedence. Boldly apply your cloud and distributed systems expertise to build deterministic tests and reusable environments.

Micro1
VerifiedSenior Platform Engineer Remote Cloud Infrastructure Expert
A remote contractor role focused on cloud platform engineering to support next-gen AI training. You’ll shape how models learn and perform by providing real-world, production-grade input across distributed systems, networking, IAM, storage, and observability. Your background in cloud infrastructure and platform ownership stands in for AI-specific experience, guiding the creation of realistic reinforcement learning environments.

Micro1
VerifiedInfrastructure Engineer
We're looking for an experienced Infrastructure Engineer to build and manage the cloud and on-prem systems that power the next wave of AI model training. In this contract role, you'll apply your deep technical know-how to help shape how AI learns and performs—no prior AI experience required, just your expertise in infrastructure.

Micro1
VerifiedCloud Architect
In this contract role, you'll put your cloud architecture expertise to work helping train next-generation AI systems. By reviewing designs, building automation, and ensuring security best practices, you'll provide high-quality real-world input that directly shapes how models learn and reason. This is a chance to apply your cloud infrastructure knowledge in a cutting-edge AI context.

Mercor
VerifiedCloud / DevOps Engineer (Infra & IaC)
A leading AI lab building foundational Large Language Models needs a Cloud/DevOps Engineer to improve the quality of its training and inference infrastructure. The role centers on hands-on work with Kubernetes, AWS services, and Infrastructure-as-Code. You will design tasks, evaluate solutions, and write technical feedback that helps models reason about cluster failures, AWS integration, and CI/CD pipelines. The position is full-time (40 hours/week), remote within the United States, and structured as a W-2 contract through Cincinnatus LLC.

Micro1
VerifiedSenior DevOps Engineer
This remote contractor role offers a chance to apply your DevOps expertise to a unique mission: training next-generation AI systems. You'll design and manage the cloud infrastructure and pipelines that enable AI models to learn and improve, using your existing skills rather than requiring prior AI experience.

