Data Engineer · I build pipelines that don't page you at 3am.
I work on the batch and streaming side of the modern data stack — designing medallion-architecture lakehouses, wiring up ingestion from messy upstream sources, and putting enough testing and observability around it that the data is actually trustworthy when it lands.
Lately that means a lot of Databricks + AWS: Delta Lake, Spark Structured Streaming, Unity Catalog governance, dbt and Great Expectations for data quality, and Terraform / Databricks Asset Bundles so none of it lives as clicked-together config.
- 🔭 Currently going deeper on streaming semantics — watermarking, CDC patterns, stream-static joins — and on making pipelines cheaper, not just correct.
- 🌱 Picking up GCP to round out the AWS/Azure side.
- 🤝 Open to collaborating on open-source data tooling, especially anything in the data-quality or orchestration space.
- 💬 Ask me about lakehouse design, Spark performance tuning, or why your CDK deploy is failing on an ARM64 mismatch.
- 📫 Reach me at set.night.bias@cloak.id
| Project | What it does | Stack |
|---|---|---|
| faker-api-test | Mock API that generates synthetic security-event data with full Snyk field parity, for volume- and load-testing downstream pipelines. Deploys natively as a Databricks App. | Node.js · Express · FakerJS · Databricks Apps |
| Databricks_RAG | Retrieval-augmented generation pipeline built on Databricks — chunking, embedding, and vector retrieval over a document corpus. | Databricks · Python · Vector Search |
| Databricks-Unity-Catalog-Project | Governed ingestion → transformation → workflow deployment with catalog-level lineage and access control end to end. | Azure Databricks · Unity Catalog · Delta Lake |
| Formula1-Azure-Project | Ingests Ergast API race data and models it into curated tables for driver and constructor analysis. | Azure Databricks · PySpark · ADF · Delta Lake |
| Kaggle_ETL | End-to-end ETL pulling Kaggle datasets through cleaning and transformation into an analysis-ready layer. | Python · Pandas · PySpark |
| Youtube_Comments_NLP | Collects YouTube comment data and runs sentiment and topic analysis over it. | Python · NLP · Jupyter |
| Web-Scraping-Google-Jobs | Scrapes and normalizes job postings into a structured, queryable dataset. | Python · BeautifulSoup · Pandas |


