Lidiya - Data Engineer | ContraWork by Lidiya
Lidiya

Lidiya

Backend & Data Engineer

Ready for work

Lidiya is ready for their next project!

Cover image for Data Trust Engine is a
Data Trust Engine is a modular data engineering platform designed for environments where the reliability of external data cannot be taken for granted. To increase confidence in incoming data, the platform cross-validates the primary data source against an independent secondary source through a dedicated Data Trust Layer. Built-in Data Reconciliation detects inconsistencies, validates data integrity, and routes suspicious records to a quarantine workflow while generating observability events and operational alerts. Rather than attempting to determine which source is objectively correct, the platform exposes inconsistencies, assigns trust signals, and provides engineers with the information required to make informed operational decisions. Automatic failover is intentionally limited to infrastructure failures. The platform switches to a secondary source only when the primary source becomes unavailable or stops responding. During normal operation, secondary sources are used exclusively for reconciliation and anomaly detection instead of replacing the primary ingestion pipeline. The platform follows a modular architecture based on Medallion Architecture, Dagster, FastAPI, and cloud-ready storage abstractions, allowing each layer to evolve independently while remaining extensible for future integrations.
0
80
DataForge MultiCloud is an end-to-end CRM data platform built around a modular ETL architecture and Medallion Design Pattern (Bronze → Silver → Gold). The pipeline extracts contact data from HubSpot, validates records with Pydantic, separates invalid data into quarantine storage, and loads clean datasets into BigQuery for transformation and analytics. Data processing is organized into dedicated layers for ingestion, validation, warehousing, orchestration, and reporting, allowing components to evolve independently. The project combines Python, Apache Airflow, dbt, BigQuery, Azure Blob Storage, DuckDB, Docker, and Pydantic to create a production-style data workflow with data quality controls, incremental processing, warehouse modeling, and analytics-ready marts. Special attention was given to operational reliability through invalid-record isolation, warehouse validation workflows, modular loaders, documented deployment procedures, and comprehensive project documentation covering architecture, data flow, warehouse layers, orchestration, testing, and deployment.
0
158
This project was designed as a modular data pipeline for automated freelance job monitoring and delivery. The architecture separates data collection, filtering, deduplication, AI-powered enrichment, storage, and notification layers, allowing each component to evolve independently. Telegram serves as the delivery channel for alerts, while the core system focuses on reliable data processing, automation, and workflow orchestration.
0
253
Tell me about your data, your sources, your workflows — bring me the chaos from databases, CRMs, and scattered APIs, and I’ll turn it into clean ETL pipelines and a real-time dashboard you’ll actually enjoy working with. While you focus on your business, the system keeps running quietly in the background 24/7. One command, one browser launch, and your live analytics are ready, just like in this demo. P.S. All data shown in this demo was synthetically generated for presentation purposes. No commercial secrets were harmed. Enjoy 💫
0
348