Analyzed approximately 600,000 column headers across a large collection of Google Sheets to identify recurring data fields, naming inconsistencies, and common schema patterns. Built an automated extraction and analysis workflow to process the spreadsheets at scale, normalize similar headers, group related fields, and surface the data structures most frequently requested across projects. The findings were used to support schema standardization and improve consistency in future data collection workflows.