Headcount: 1 | Work Arrangement: Remote | Employment Type: Direct Hire | Role Type: Newly Created Position
Our client is a growing SaaS company in the damage prevention, asset integrity, and infrastructure-protection space, serving energy, utility, telecom, and infrastructure companies across North America.
Role Overview
Our client is building a modern, multi-cloud, enterprise-grade data estate — a unified Databricks-based platform, larger in scope than a standard lakehouse, that centralizes data across the company's products and cloud environments (AWS, Azure, and GCP) and feeds multiple product lines, primarily the damage prevention platform. This role will primarily support that platform by integrating data from multiple sources, with occasional work touching other product lines. This is a skill-first role: strong fundamentals in statistics, machine learning, and deep learning, combined with genuinely practical GenAI/agentic-systems experience (LangChain, agent workflows, MCPs), matter more than direct industry-domain background.
Key Responsibilities
- Contribute to medallion architecture pipelines (Bronze → Silver → Gold) using Databricks, and support column-level lineage and governance initiatives (targeting at least 95% lineage coverage).
- Explore, prototype, evaluate, and productionize machine learning and GenAI solutions across use cases including forecasting, anomaly detection, NLP, Retrieval-Augmented Generation (RAG), LLM-powered assistants/copilots, and predictive analytics.
- Package and manage models using Unity Catalog model management/registries, and design batch and streaming inference architectures where appropriate.
- Partner with Product and business stakeholders to define success metrics, KPIs, and A/B testing strategies, and move successful experiments from prototype to production with clear SLAs, monitoring, and documentation.
- Build production workflows, jobs, and notebooks as infrastructure-as-code using Databricks Asset Bundles (DABs), with CI/CD via GitHub Actions.
- Contribute business metrics, definitions, and semantic models to Unity Catalog, supporting consumption through Power BI and Databricks AI/BI.
- Implement secure data and ML architectures using RBAC/ABAC within Unity Catalog, and support compliance requirements (SOC 2, ISO 27001, GDPR, PIPEDA).
Requirements
- 3–6 years of experience in Data Science, Machine Learning, or ML Engineering, with a track record of taking models from development through production.
- Strong basic understanding of statistics, machine learning, and deep learning fundamentals.
- Practical, hands-on GenAI/LLM experience: prompt engineering, RAG, vector databases/vector stores, LLM evaluation, agentic workflows, LangChain, and MCPs.
- Strong programming and data skills in Python, SQL, and Spark/PySpark, with hands-on Databricks experience (Delta Lake, Unity Catalog, DBSQL, Jobs and Workflows, medallion architecture).
- Bachelor's degree is sufficient — a master's in statistics is a plus, as is a well-regarded certification (e.g. DeepLearning.AI, Coursera courses by instructors such as Andrew Ng).
- Strong communication and collaboration skills.
Nice-to-Have
- Experience with Microsoft Azure, AWS, or Google Cloud Platform (GCP).
- Experience with geospatial data and analytics (PostGIS, spatial joins, spatial indexing, GIS-based feature engineering).
- Experience with streaming and real-time data, including Structured Streaming and Change Data Capture (CDC).
- Hands-on experience with MLflow and Unity Catalog Model Serving.
- Experience working in utilities, energy, infrastructure, or public works industries.
Benefits
- Be part of a dynamic and growing company that is well-respected in its industry.
- Competitive compensation based on experience and qualifications.
- Health Insurance coverage.
As part of our hiring process, this role may use artificial intelligence or automated tools to assist with reviewing and screening applications. These tools support, but do not replace, human judgment in making hiring decisions.
Your application will only be counted once you complete the full registration process on the KeyStone platform, including creating your profile, uploading your CV, and submitting your application. The AI interview is optional and encouraged, but is not required for your application to be counted.
* Questions marked with an asterisk are required for eligibility.