Shrief Khamis - Data Engineer | Contra
Work by Shrief Khamis
Sign Up
Post a job
Sign Up
Log In
Shrief Khamis
Backend Engineer | APIs, Data Pipelines & AI
Message
Follow
New to Contra
Shrief is ready for their next project!
Cairo, Egypt
Work
Services
About
Cairo, Egypt
0
Designed and built a configurable Python engine for generating large, realistic relational datasets with preserved primary-key and foreign-key relationships. The system supports reusable business templates, weighted distributions, multiple ID strategies, and scalable generation across related tables. Performance work increased throughput to more than 300,000 rows per second in large benchmark runs. The project is being extended with dirty-data injection, validation reports, and additional output formats.
0
8
0
Analyzed approximately 600,000 column headers across a large collection of Google Sheets to identify recurring data fields, naming inconsistencies, and common schema patterns. Built an automated extraction and analysis workflow to process the spreadsheets at scale, normalize similar headers, group related fields, and surface the data structures most frequently requested across projects. The findings were used to support schema standardization and improve consistency in future data collection workflows.
0
13
0
Designed and built an AI-assisted research system that automates the process of locating, prioritizing, and extracting information from websites. The system crawls relevant pages, retrieves the most useful content, and uses Retrieval-Augmented Generation (RAG) to answer user questions with supporting quotations and source references. Built with a modular architecture to support multiple research tasks while minimizing unnecessary crawling and reducing manual research effort.
0
17
0
Designed and built an end-to-end data collection pipeline used by a distributed team. A custom browser extension captured structured data during normal workflows and sent it to a FastAPI service for validation, transformation, and centralized processing before storage in PostgreSQL on AWS RDS. The system replaced a manual collection step, handled roughly 5,000 records per day, and was designed for reliable API communication, maintainability, and future expansion.
0
21