Data Engineer (Databricks)
What is the role?
As a Data Engineer at Tensorplay, you’ll build the pipelines that turn raw enterprise data into reliable, AI-ready datasets on the Databricks lakehouse. You’ll work alongside our data architects and ML engineers to ingest, transform, and serve data for dashboards, models, and GenAI applications — and make sure it stays accurate, fresh, and cost-efficient in production.
About you
You enjoy writing clean, well-tested data code and you’ve shipped Spark pipelines that run reliably every day. You’re comfortable in Databricks notebooks and jobs, you understand how Delta Lake works under the hood, and you take data quality seriously. You like seeing how your work gets used downstream — especially when it powers ML and AI.
Your Responsibilities
- Build and maintain batch and streaming pipelines on Databricks using PySpark and Spark SQL
- Implement medallion (bronze/silver/gold) layers with Delta Lake, Auto Loader, and Delta Live Tables / Lakeflow
- Orchestrate and schedule pipelines with Databricks Workflows and manage them through CI/CD
- Apply data quality checks, expectations, and monitoring to keep pipelines trustworthy
- Manage tables, permissions, and lineage in Unity Catalog
- Tune Spark jobs and Delta tables for performance and cost (partitioning, Z-ordering, liquid clustering)
- Prepare datasets and document pipelines for ML training, feature tables, and RAG applications
Requirements
- 3+ years of data engineering experience, with at least 1–2 years on Databricks
- Strong Python and SQL skills, with hands-on PySpark experience
- Working knowledge of Delta Lake and lakehouse concepts
- Experience with at least one cloud platform (Azure, AWS, or GCP) and its storage services
- Familiarity with Git, code review, and CI/CD practices for data pipelines
- Solid understanding of data modelling and ETL/ELT design
Nice to Have
- Databricks Certified Data Engineer Associate or Professional certification
- Experience with Structured Streaming, Kafka, or Event Hubs
- Exposure to MLflow, feature stores, or vector search
- Experience with dbt, Airflow, or Terraform
We Offer
- Competitive salary with performance-based bonuses
- Fully remote with flexible working hours
- Databricks certification and learning budget
- Hands-on work across varied enterprise data and AI projects
- Mentorship from senior data architects and a clear growth path