Site Reliability Engineer | $20-$70/hr Remote
Overview
As a Site Reliability Engineer at micro1, you'll put your infrastructure and operations expertise to work in an unexpected context: training the next wave of AI systems. Your responsibility is to build and maintain rock-solid platforms using Linux, Kubernetes, and Prometheus — ensuring models learn from reliable, real-world data. No AI background required; your SRE skills are the foundation.
What You'll Do7
- 1Architect, deploy, and maintain scalable infrastructure powered by Linux clusters, Kubernetes orchestration, and Prometheus monitoring.
- 2Continuously observe system health, analyze performance metrics, and spot potential bottlenecks or failure points before they escalate.
- 3Automate operational workflows — from provisioning to incident response — to reduce manual toil and improve system resilience.
- 4Lead incident response efforts, perform root cause analysis, and refine on-call runbooks to shorten resolution times.
- 5Work closely with development and operations teams to orchestrate smooth deployments and maintain high availability across services.
- 6Write clear, actionable documentation and runbooks so the whole team can operate efficiently and share knowledge.
- 7Promote SRE best practices, security standards, and compliance measures throughout the customer's ecosystem.
Requirements7
- 1Expert-level command of Linux system administration, including troubleshooting, tuning, and performance analysis.
- 2Advanced hands-on experience with Kubernetes — cluster deployment, lifecycle management, and operational troubleshooting.
- 3Deep proficiency with Prometheus for metrics collection, alerting, and building dashboards that inform decision-making.
- 4Strong scripting skills in Bash, Python, or a similar language for automation and tooling.
- 5Excellent written and verbal communication skills, with a track record of documenting systems and processes clearly.
- 6Proven experience as a Site Reliability Engineer or similar role in high-availability, production-critical environments.
- 7Demonstrated ability to solve problems proactively and collaborate effectively across teams.
Who Should Apply
We're looking for an experienced SRE who takes pride in building systems that never go down — and who enjoys the challenge of applying those skills to help train AI. You're comfortable automating everything in sight, diving into incidents headfirst, and sharing your knowledge through documentation. A background in high-growth or large-scale environments is a plus, but your hands-on expertise with Linux, Kubernetes, and Prometheus is what really matters.
Salary Insight
This is a remote contract position paying $20 to $70 per hour, depending on experience and qualifications.
Required Skills
Application Tip
In your application, highlight two specific incidents where you used Kubernetes and Prometheus to diagnose and resolve a production issue. Walk through the problem, your monitoring approach, and the automation you built afterward to prevent recurrence.
Similar open positions
Explore active roles that match your skills and interests.
Micro1
VerifiedDevOps Engineer | $20-$70/hr Remote
This DevOps Engineer role is your opportunity to apply your cloud infrastructure expertise to directly influence how next-generation AI systems learn and perform. You'll design and automate scalable environments using AWS and Kubernetes, ensuring reliability and efficiency for model training workflows. No prior AI experience is needed—your proven DevOps skills are what matter most.
Micro1
VerifiedAI Engineer | $30-$130/hr Remote
As an AI Engineer at micro1, you'll play a hands-on role in shaping how next-generation AI systems learn, reason, and perform. Your domain expertise — even without a formal AI background — will be the key ingredient for training models on real-world, high-quality data. Bring your strong engineering and cloud skills, and we'll provide the platform to deploy and scale impactful machine learning solutions.
Micro1
VerifiedInfrastructure Engineer | $20-$70/hr Remote
We're looking for an experienced Infrastructure Engineer to build and manage the cloud and on-prem systems that power the next wave of AI model training. In this contract role, you'll apply your deep technical know-how to help shape how AI learns and performs—no prior AI experience required, just your expertise in infrastructure.
Micro1
VerifiedSenior DevOps Engineer | $30-$130/hr Remote
This remote contractor role offers a chance to apply your DevOps expertise to a unique mission: training next-generation AI systems. You'll design and manage the cloud infrastructure and pipelines that enable AI models to learn and improve, using your existing skills rather than requiring prior AI experience.
Micro1
VerifiedDevOps Engineer | $30-$80/hr Remote
This remote contract role blends DevOps engineering with a unique mission — helping train next-generation AI systems. You'll use your infrastructure expertise to build resilient, scalable cloud environments that power AI learning. No prior AI experience is needed; what matters is your ability to design and manage complex systems across AWS, Azure, and GCP.
Micro1
VerifiedRunOps Support – Platform and Infra | $25-$50/hr Remote
This role puts your infrastructure expertise to work powering next-generation AI systems. You'll monitor and maintain production environments, respond to incidents, and improve platform reliability — all while helping train smarter AI models. No prior AI experience needed; your hands-on knowledge of Linux, containers, and cloud platforms is what counts.