Projects using Databricks in Columbus
Projects using Databricks in Columbus
Sign Up
Post a job
Sign Up
Log In
Filters
2
Projects
People
Message
1
VJ Sojitra
I engineered a large-scale package-network intelligence pipeline using Azure Databricks and PySpark. The system joined approximately 4.5 billion historical records and transformed them into reusable fingerprint signals covering package movement and network behavior. I developed time-windowed feature logic using native Spark expressions rather than Python UDFs, preserving distributed performance at scale. The pipeline produced an approximately 200-million-row fingerprint dataset and a 22-million-row operational output that classified package activity along a cold-to-hot temperature scale. The resulting signals supported downstream monitoring, prioritization, and operational analysis. The architecture shown here is a sanitized representation; client identifiers, proprietary schemas, infrastructure paths, and business rules have been removed.
1
157
Explore projects