AI Operations Engineer
Overview
We're seeking a hands-on engineer to maintain and optimize our GPU-powered AI inference platform that serves large language models to internal applications. You'll own the full stack from bare-metal OS setup through production deployment, monitoring, and capacity planning to keep the infrastructure reliable and scalable. This role combines deep systems engineering with AI/ML operations, requiring a mix of automation skills and performance tuning.
What You'll Do6
- 1Provision and automate Linux servers: harden the OS, install GPU drivers and toolkits, set up Python environments, and deploy LLM inference engines using version-controlled automation (e.g., Ansible) for repeatable, auditable deployments.
- 2Deploy and manage LLM inference engines (like vLLM or TensorRT-LLM) across a fleet of enterprise GPUs, handling multi-GPU model sharding, quantization strategies, and GPU memory/cache sizing to maximize concurrency and throughput.
- 3Operate the API gateway layer for load balancing, model routing, API key management, and token counting, along with supporting services such as database backends, caching layers, and reverse proxies with proper access controls.
- 4Instrument and maintain the observability stack: collect GPU-level telemetry (utilization, memory, temperature), track key metrics (latency percentiles, tokens per second, concurrent users), and set alerting thresholds for capacity and hardware health.
- 5Collaborate closely with application development teams to onboard new workloads, review and optimize system prompts, and troubleshoot model behavior – distinguishing between model limitations, prompt issues, and infrastructure problems.
- 6Support model fine-tuning workflows by deploying LoRA/QLoRA adapters into the serving infrastructure and understanding how fine-tuned models differ from base models in terms of serving requirements.
Requirements6
- 1At least 4 years of hands-on Linux systems engineering experience, with no fewer than 2 years involving GPU infrastructure or ML/AI workloads in a production environment.
- 2Proven track record of deploying and operating LLM inference engines in production, including experience with model quantization, multi-GPU parallelism, and performance tuning.
- 3Strong working knowledge of the NVIDIA GPU software stack (drivers, toolkits, runtime libraries) and the ability to diagnose common failure modes and compatibility issues.
- 4Proficiency with infrastructure-as-code and configuration management tools (especially Ansible), plus solid understanding of Linux containerization (Docker, rootless operation, service management integration).
- 5Comfortable scripting in Python and shell for operational tooling, and disciplined use of Git for version control (branching, tagging, commit practices).
- 6Excellent troubleshooting and communication skills – able to document procedures, write runbooks, and explain infrastructure constraints to development teams, especially in environments with restricted internet access.
Who Should Apply
This role is perfect for a systems engineer who thrives on keeping AI infrastructure running smoothly and loves getting their hands dirty with GPU clusters. You're methodical, can work independently, and communicate clearly with developers about platform capabilities and constraints. If you enjoy optimizing performance, troubleshooting at the kernel level, and staying ahead of new model releases, this is for you.
Required Skills
Application Tip
When applying, include a specific example of a time you optimized LLM serving throughput or resolved a tricky GPU driver issue in a production environment. Highlight your experience with Ansible or similar IaC tools and any work with air-gapped deployments.
Similar open positions
Explore active roles that match your skills and interests.
Optim Hire
VerifiedSenior AI ML Engineer - GenAI LLM Data Science and MLOps
We're seeking a Senior AI/ML Engineer with 6-7 years of experience to lead the design and deployment of scalable AI solutions. This on-site role focuses on Generative AI, Large Language Models (LLMs), Data Science, and MLOps — building production-ready systems that power real-world products. You'll collaborate with cross-functional teams, mentor junior engineers, and shape our AI strategy.
Micro1
VerifiedForward Deployed Engineer | $300K - $650K/yr | Remote
This role sits at the intersection of applied AI, ML infrastructure, and partner-facing product development. You'll work directly with leading AI labs and enterprises to transform ambiguous research questions into production-grade systems. As a Member of Technical Staff, you'll own everything from data curation and LLM agent workflows to deployment and partner success — all while operating in a fast-moving, remote-first environment.
Talentxo
VerifiedAl/ML Engineer
We are looking for a hands-on AI/ML Engineer to build and deploy intelligent AI solutions across a communication platform. The role involves developing LLM-powered applications, voice AI capabilities, NLP pipelines, and production-grade ML systems that enable automation and real-time decision-making.
Optim Hire
VerifiedSenior AIML Engineer LLM RAG Computer Vision MLOps
We're looking for a senior engineer who can build and deploy production-ready AI systems that blend large language models, retrieval-augmented generation, computer vision, and robust MLOps. In this role, you'll create end-to-end pipelines that turn messy documents and video streams into structured insights, then keep those systems running smoothly at scale. It's a hands-on position that demands both strong machine learning fundamentals and the software engineering chops to ship reliable, high-performance solutions.
Talentxo
VerifiedAI Architect (GenAI, Python)
Design, develop, and deploy scalable AI solutions with expertise in machine learning, deep learning, and Generative AI. Architect end-to-end AI systems using LLMs, prompt engineering, and fine-tuning, while ensuring robust MLOps practices.
Optim Hire
VerifiedSenior AI ML Engineer
VectorStack is looking for a seasoned AI/ML engineer to build and scale AI-powered backend systems. You'll lead the integration of machine learning models into production, design robust APIs, and collaborate across teams. This on-site role in Bengaluru is ideal for engineers with 5+ years of experience who thrive on solving complex problems and mentoring others.