Data Engineer (Mid-Level)
Location: India | Remote | Direct Hire | Newly Created Position | Headcount: 1
About the Company
Our client is a growing SaaS company that provides cloud-based software for damage prevention, asset integrity, stakeholder engagement, and land management, helping energy, utility, telecom, and infrastructure companies protect their critical network infrastructure. With nearly three decades of industry experience, they continue to expand their platform with new data-driven and AI-powered capabilities.
Role Overview
Our client is building a modern, multi-cloud, enterprise-grade data estate — a unified Databricks-based platform, larger in scope than a standard lakehouse, that centralizes data across their products and cloud environments (AWS, Azure, and GCP). This is a hands-on implementation role, working closely with the Senior Data Architect to translate the platform's architecture and technical standards into reliable, production-ready pipelines. This is a skill-first role: deep, hands-on Databricks and medallion architecture expertise is the primary requirement. Change Data Capture (CDC) experience is a high priority. Machine learning exposure is a nice-to-have, secondary to the core data engineering skill set. Effective, efficient use of GenAI tools in day-to-day work is expected. A bachelor's/undergraduate degree is sufficient, and a Databricks certification is a strong plus.
Key Responsibilities
Build, maintain, and enhance data ingestion pipelines across AWS, Azure, and GCP, following established architecture and engineering patterns.
Develop both batch and streaming pipelines using Databricks Workflows, Apache Spark/PySpark, SQL, Delta Live Tables, and Databricks Lakeflow components.
Implement Bronze → Silver → Gold medallion architecture patterns for ingestion, transformation, cleansing, and standardization.
Implement Change Data Capture (CDC) and Slowly Changing Dimensions (SCD Type 1 and Type 2).
Handle schema evolution and changing source-system structures.
Implement data validation, reconciliation, and quality rules as part of pipeline processing.
Configure and maintain Delta Lake storage structures, tables, schemas, partitions, and optimization routines (OPTIMIZE, Z-ORDER, VACUUM).
Assist with implementation of metadata, cataloging, and lineage standards using Unity Catalog.
Implement automated data-quality checks, profiling, validation, and monitoring in accordance with enterprise governance standards.
Build, schedule, monitor, and maintain production workflows using Databricks Workflows, Delta Live Tables, and Azure Data Factory (ADF).
Contribute to CI/CD pipelines, including source control, automated testing, and DEV → QA → PROD promotion.
Monitor production pipelines, troubleshoot failed jobs, and support pipeline recovery.
Work directly with the Senior Data Architect to translate architecture designs into actionable implementation tasks.
Document data pipelines, data flows, data dictionaries, transformation logic, data-quality rules, and operational procedures.
Requirements
Hands-on experience implementing medallion architecture within a Databricks environment.
3–5 years of experience in Data Engineering, ETL development, or cloud data platform engineering.
Strong programming and data skills in Python, SQL, and Apache Spark/PySpark.
Change Data Capture (CDC) experience.
Experience with at least one major cloud platform — Microsoft Azure preferred.
Understanding of core data-engineering concepts: data modeling, data quality, schema evolution, and pipeline monitoring/troubleshooting.
Basic understanding of data-security practices (RBAC, encryption, credential/secret management).
Effective, efficient use of GenAI tools in day-to-day engineering work.
Bachelor's/undergraduate degree — no advanced degree required.
Nice-to-Have
Databricks certification (strong plus) or Microsoft Azure certification (e.g. DP-203).
Experience with metadata/governance platforms such as Unity Catalog, Microsoft Purview, or AWS Glue Data Catalog.
Experience with orchestration tools such as Azure Data Factory (ADF), Databricks Workflows, or Apache Airflow.
Machine learning exposure.
Experience with Git-based development, CI/CD, and DevOps practices.
Geospatial/GIS data experience, or BI semantic layers (particularly Power BI).
Benefits
Be part of a dynamic and growing company that is well-respected in its industry.
Competitive compensation based on experience and qualifications.
Health Insurance coverage.
As part of our hiring process, this role may use artificial intelligence or automated tools to assist with reviewing and screening applications. These tools support, but do not replace, human judgment in making hiring decisions.
Your application will only be counted once you complete the full registration process on the KeyStone platform, including creating your profile, uploading your CV, and submitting your application. The AI interview is optional and encouraged, but is not required for your application to be counted.
* Questions marked with an asterisk are required for eligibility.