About the role
As a Senior Data Engineer, you will work across Google Cloud Platform (GCP) and Databricks, building and evolving scalable data pipelines, data products, and analytics-ready datasets.
The current environment is approximately 60–70% GCP-focused, so strong hands-on GCP experience is mandatory. Initially, your work will primarily focus on the existing GCP platform, with increasing ownership of Databricks pipelines, integrations, and lakehouse capabilities as the platform evolves.
You will also work closely with Data Scientists and Analytics teams to deliver trusted, governed, AI-ready and reporting-ready datasets.
Key Responsibilities
1) GCP Data Engineering
Design and maintain scalable data architectures and pipelines on GCP.
Build and optimize solutions using BigQuery, Dataflow, Cloud Run, Composer, and GCS.
Develop reliable pipelines supporting analytics, reporting, and machine learning workloads.
Translate business requirements into scalable technical solutions.
Maintain high standards around performance, reliability, security, and cost.
2) Databricks Engineering
Own and develop Databricks pipelines and integrations.
Work with Delta Lake, Unity Catalog, Workflows, MLflow, and Spark.
Support the gradual expansion of Databricks within the wider data platform.
Ensure GCP and Databricks workloads follow consistent engineering, governance, and security standards.
Help shape future lakehouse architecture and integration patterns.
3) Data Transformation & Semantic Layer
Develop transformation workflows using SQL, Python, and dbt.
Build and maintain reusable semantic layers and data models.
Deliver datasets that are ready for analytics, reporting, AI, and ML use cases.
Establish standards for testing, documentation, versioning, and deployment.
4) Data Quality, Governance & Reliability
Implement data quality controls and automated testing.
Ensure data accuracy, governance, security, and accessibility.
Monitor pipeline health, freshness, performance, and operational stability.
Troubleshoot incidents and drive continuous improvement.
Support data lineage, access controls, and lifecycle management.
5) AI / ML Enablement
Work closely with Data Scientists and Analytics teams.
Understand ML workflows and the data requirements behind them.
Build trusted and reusable datasets for ML and AI use cases.
Support ML pipelines and integrations, including the use of MLflow where applicable.
Help create AI-ready datasets through the semantic layer.
6) Platform Ownership & Engineering Standards
Contribute to infrastructure and deployment standards using Terraform, Docker, CI/CD, and Git.
Promote engineering best practices around testing, documentation, code reviews, and operational ownership.
Provide technical guidance and mentor other engineers.
Document architectures, data flows, integrations, and operational processes.
7) Collaboration & Stakeholder Management
Work with Product Owners, Data Scientists, Analysts, Engineers, and business stakeholders.
Translate business needs into scalable technical solutions.
Communicate technical decisions, trade-offs, risks, and progress clearly.
Take ownership of solutions from design through production.
Strong hands-on Google Cloud Platform (GCP) experience is mandatory.
Strong experience with BigQuery and GCP-based data pipelines.
Nice-to-Have
Experience with:
Delta Lake
Unity Catalog
Databricks Workflows
MLflow
Apache Spark
GCP services such as:
Dataflow
Cloud Run
Composer / Airflow
Google Cloud Storage
Perks & Benefits
We believe taking care of our people is non-negotiable. Here's what being a Star means:
"Gapstars is committed to a diverse and inclusive workplace. We are an equal-opportunity employer and do not discriminate based on race, national origin, gender, disability, or age. Your personal information collected during the application process is handled following our privacy policy and used exclusively for recruitment and hiring purposes."