Projects using KafkaProjects using KafkaProblem:
Many organizations still process invoices manually by reading PDF documents and entering key details (invoice number, vendor, amount, etc.) into systems. This process is slow, error-prone, and difficult to scale, and it also makes it harder to detect duplicate invoices or incorrect totals.
Solution:
This project builds an automated invoice processing pipeline that converts uploaded invoice PDFs into structured data. It uses OCR to extract text, LLMs to identify invoice fields, validation checks to ensure correctness, and Kafka-based event streaming to manage the processing pipeline. The extracted data is stored in PostgreSQL and visualized through a dashboard, enabling faster, scalable, and more reliable invoice processing. Problem:
Batch data processing was not suitable for real-time analytics and scalable cloud-based data ingestion.
Solution:
Created a real-time streaming pipeline using Kafka on AWS EC2, stored processed data in S3, cataloged it with AWS Glue, and queried it with Amazon Athena.
Tools:
Python, Apache Kafka, AWS EC2, Amazon S3, AWS Glue, Amazon Athena, Pandas
Result:
Built an end-to-end cloud data streaming workflow that supports real-time ingestion, storage, cataloging, and SQL-based analytics.