Lead Data Engineer
Overview
We're looking for a Lead Data Engineer to design and optimize enterprise-grade data pipelines on Google Cloud Platform. This role centers on building scalable, real-time and batch processing solutions using Apache Beam, Dataflow, Java, and BigQuery. You'll collaborate closely with architects, analysts, and business stakeholders to ensure data quality, performance, and cost efficiency.
What You'll Do8
- 1Architect and implement scalable ETL/ELT pipelines using GCP services like Dataflow and BigQuery.
- 2Develop and maintain both real-time and batch data processing workflows with Apache Beam and Dataflow.
- 3Write robust backend processing logic in Java to support data transformations and integrations.
- 4Work extensively with BigQuery for data warehousing, analytics, and performance tuning.
- 5Integrate data from diverse sources including APIs, databases, and streaming platforms (e.g., Pub/Sub, Kafka).
- 6Optimize pipeline performance, cost, scalability, and reliability across the GCP stack.
- 7Collaborate with cross-functional teams (DevOps, Architects, Analysts) to enforce data governance, security, and monitoring best practices.
- 8Troubleshoot production issues and deliver long-term, scalable solutions.
Requirements8
- 17+ years of hands-on experience in data engineering with a strong focus on Google Cloud Platform services.
- 2Deep expertise in Apache Beam and Google Dataflow for building both batch and streaming pipelines.
- 3Strong programming skills in Java and experience working with BigQuery for data warehousing and analytics.
- 4Proficiency in SQL, data modeling, and ETL/ELT best practices.
- 5Familiarity with CI/CD pipelines, Git, and Agile development methodologies.
- 6Experience with additional GCP tools like Pub/Sub, Cloud Composer, Cloud Storage, or Dataproc is a plus.
- 7Exposure to Python or Spark is desirable.
- 8A Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field is preferred; GCP certifications are an added advantage.
Who Should Apply
This role is ideal for a seasoned data engineer who thrives on building and optimizing large-scale data pipelines in the cloud. You should have a strong command of GCP services, especially Dataflow and BigQuery, and be comfortable writing production-quality Java code. If you enjoy collaborating across teams, solving complex data challenges, and driving performance improvements, you'll fit right in.
Required Skills
Application Tip
When applying, include a brief case study of a specific pipeline optimization you led—mention the tools (e.g., Dataflow, BigQuery) and the impact on performance or cost. This demonstrates your hands-on expertise and problem-solving mindset.
Similar open positions
Explore active roles that match your skills and interests.
Talentxo
VerifiedData Engineer (Databricks, Python)
This role focuses on architecting and implementing enterprise-grade Lakehouse solutions using Databricks and Apache Spark. The position involves building scalable batch and real-time data pipelines, leading ML lifecycle management, and managing cloud-native deployments on Microsoft Azure. It requires 10-15 years of experience with deep expertise in Databricks and Data Engineering.
Optim Hire
VerifiedProject Lead Data Lakehouse Platform
We're seeking an experienced project lead to guide the PMO function for building and deploying a robust Data Lakehouse platform. In this role, you'll serve as the key liaison between the client, system integrator, and technology teams, ensuring the platform aligns with business goals, security standards, and regulatory requirements—especially within the banking sector. Your focus will be on coordinating delivery, tracking progress, and keeping everyone aligned toward a successful rollout.
Optim Hire
VerifiedSenior BackEnd Engineer
Tapistro is harnessing AI to transform B2B go-to-market orchestration, shifting from scattered campaigns to targeted precision. As a founding team member, you'll build scalable backend systems that power next-gen AI solutions, directly shaping the company's trajectory and the broader AI landscape.
Optim Hire
VerifiedSenior Data Scientist
We're looking for a seasoned Data Scientist who can take machine learning models from idea to production. You'll work across the full stack—building, training, and deploying Generative AI and LLM solutions on GCP/Azure/AWS, backed by solid engineering practices like CI/CD and API design. This hybrid role is based in Chennai, Bangalore, Pune, or Coimbatore.
Optim Hire
VerifiedData Scientist
We are looking for a seasoned Data Scientist to lead the design and deployment of machine learning models within a modern Python-based cloud architecture. You'll take ownership of the full ML lifecycle, from data ingestion to model scoring, while guiding a team and collaborating with executives. This role combines deep technical expertise with strategic oversight to drive AI initiatives forward.
Optim Hire
VerifiedAI Lead
We're looking for an experienced AI Lead to drive our AI strategy, architecture, and delivery. You'll oversee end-to-end AI initiatives, mentor a team of data scientists and engineers, and collaborate closely with business and engineering leaders to build production-grade AI systems. This role is ideal for someone who enjoys shaping AI roadmaps and delivering measurable business impact.