United States
about 2 hoursData Engineer – Insights (AI/ML)
About Our Client
Our client is a leading SaaS company with nearly three decades of experience helping energy, utility, telecom, and infrastructure companies protect their critical network infrastructure. They are now expanding their platform with new data-driven and AI-powered capabilities.
About the Role
Our client is building a new AI-driven threat and risk management platform for pipeline asset integrity. It brings together a governed, cross-product data platform built on Databricks and Azure, an AI-powered ingestion layer that normalizes, repairs, and enriches customer data, and a reusable analytical layer that runs risk models against that data. As a Data Engineer, you will build the pipelines that make this platform real. This is a new project, so you will help build it from the ground up, implementing patterns set by the Data Architect and working closely with Data Scientists on AI-assisted data ingestion. The team is US-based and works daily with colleagues in India.
Location: Remote (United States)
Schedule: Monday-Friday, with working hours that overlap with the India team, roughly 8:00-11:00 AM Eastern Time.
Employment Type: Full-time
Key Responsibilities
- Build and maintain ingestion pipelines for structured and semi-structured sources, including GIS, inline inspection data, SCADA, maintenance systems, and enterprise systems of record.
- Implement batch and streaming ingestion using Databricks Workflows, Spark, PySpark, SQL, and declarative pipeline tooling.
- Apply medallion architecture patterns (Bronze, Silver, Gold) for transformation, standardization, and enrichment.
- Implement change data capture (CDC), slowly changing dimensions (SCD), schema evolution, and data-validation rules.
- Normalize third-party and public data feeds (weather, soil, satellite-derived, and one-call ticket data) into the shared data model.
- Build pipelines that automate normalization of units, schemas, and semantics across inconsistent customer data, including AI-assisted mapping of new customer schemas to the target schema.
- Implement automated data-quality repair workflows, with clear provenance for every synthesized value.
- Work with Data Scientists to productionize document-extraction pipelines, and build human-in-the-loop review workflows for low-confidence extractions.
- Configure and manage Delta Lake tables, partitioning, and optimization routines, and implement lineage and cataloging with Unity Catalog.
- Build connectors to customer systems of record and support geospatial processing, including spatial joins.
- Implement data-quality tests, profiling, drift monitoring, access-control policies, and lineage for regulatory traceability.
- Build, schedule, and monitor data workflows, and own alerting and failure handling for your pipelines.
- Implement CI/CD for pipeline code, including version control, automated testing, and environment promotion.
- Troubleshoot production incidents, recover failed runs, and optimize performance and cost.
- Participate in architecture, design, and code reviews, and document pipelines, data dictionaries, and runbooks.
Requirements
- Hands-on experience building production pipelines on Databricks, including Spark/PySpark.
- Hands-on experience with Delta Lake, medallion architecture, and lakehouse engineering.
- Strong SQL skills.
- 3-5 years of experience in data engineering, ETL development, or cloud data platform engineering.
- Hands-on CI/CD experience for data pipelines.
- Experience using AI-assisted coding tools (e.g. Cursor, GitHub Copilot, Claude Code) in your development work.
- Git-based development in a code-reviewed team (GitHub, GitLab, or Bitbucket).
- Experience with at least one major cloud platform (Azure preferred).
- Familiarity with data modeling, data quality, schema evolution, and pipeline troubleshooting.
- Experience with workflow orchestration and scheduling frameworks.
- Understanding of core data-security practices, including access control, encryption, and credential management.
- Availability to work hours that overlap with 8:00-11:00 AM Eastern Time.
Nice-to-Have
- Microsoft Azure experience.
- Unity Catalog, Microsoft Purview, or comparable metadata and lineage tooling.
- Building pipelines that ingest unstructured or semi-structured documents.
- Geospatial data processing and common GIS data formats.
- Infrastructure as code (IaC) for data workloads.
- Preparing data for machine learning or probabilistic model consumption.
- Cloud or Databricks certifications.
- Experience with oil and gas or utility asset data, asset integrity concepts, or related regulatory reporting.
- Migrating customers from legacy or spreadsheet-based systems to modern data platforms.
Benefits
- Competitive salary
- Medical, dental, and vision insurance
- 401(k) plan with company match
- Generous paid time off (PTO)
- Company-paid holidays
- Remote work
As part of our hiring process, this role may use artificial intelligence or automated tools to assist with reviewing and screening applications. These tools support, but do not replace, human judgment in making hiring decisions.
Your application will only be counted once you complete the full registration process on the KeyStone platform, including creating your profile, uploading your CV, and submitting your application. The AI interview is optional and encouraged, but is not required for your application to be counted.
* Questions marked with an asterisk are required for eligibility.