Capwell's platform gives members access to a large database of professional awards. That database had grown to more than 500,000 records, but keeping it current was becoming a project of its own.
Each award could have a different source page, structure and update cycle. A record might include the award name, opening date, eligibility, entry fee, enrollment instructions, judging date, deadline extensions, rules and other details. When a source changed, someone had to find the change and decide whether the stored record should be updated.
Before this system, Capwell relied on outsourced researchers to open award URLs and compare the source information with the database. A broad refresh could take roughly six months.
I joined the project as an AI Developer and Backend Developer. My job was to build the internal AI and backend workflow that could do the first research and verification pass at scale, while keeping people involved wherever the result was unclear.
The problem was not extraction alone
Reading a page and returning a few fields was only the first step. The system also had to understand what had changed, normalize different ways of expressing the same information and avoid treating uncertain output as fact.
A useful refresh therefore needed to answer several questions for every record:
Is the original award page still available?
Which details can be extracted with enough confidence?
Does the new information actually differ from the stored record?
Is the difference a real update, a formatting change or an ambiguous case?
Can the database be updated automatically, or should a person review it?
At this volume, the workflow also had to coordinate long-running tasks without turning the entire refresh into one fragile job.
What I built
I built the AI and backend automation workflow end to end using Python, LangChain and LangGraph.
The pipeline covered:
Scheduling and orchestrating work across the award dataset
Acquiring information from each award source page
Extracting and normalizing the required fields
Comparing the new result with the existing database record
Verifying proposed changes before they reached the database
Updating confident results and routing uncertain cases to human review
Deploying and monitoring the long-running process
An individual task could take around ten minutes, so the system ran many independent tasks in parallel rather than waiting for one record to finish before starting the next.
This was an internal, terminal-run pipeline. I did not build Capwell's public website or a customer-facing dashboard. My ownership was the AI and backend system behind the refresh process.
Keeping people in the quality loop
The goal was not to remove judgment from the process. It was to stop spending human time on every record.
I added an AI oversight layer that checked the proposed results and flagged ambiguity or uncertainty. Confident matches could continue through the update path. Questionable records were separated for human review.
That changed the role of the research team. Instead of opening every URL and comparing every field manually, reviewers could focus on the cases where the automated pass found something unclear.
No extraction or language model was treated as automatically correct. The review path was part of the system, not an afterthought.
The result
The database contained more than 500,000 award records. For a comparable broad refresh scope, the new workflow completed the work in approximately three weeks, compared with roughly six months under the earlier manual process.
The practical outcomes were:
500K+ total award records handled by the workflow
A comparable refresh reduced from about six months to about three weeks
Parallel processing for tasks that could take around ten minutes each
Automated extraction, normalization, comparison and verification
Human review concentrated on ambiguous and uncertain records
One coordinated AI and backend pipeline instead of a record-by-record manual pass
I am deliberately not attaching an accuracy percentage or a fixed labor-reduction figure to the result. Those numbers were not measured in a way I can defend publicly. The cycle-time change and the database scale are the clearest evidence of the system's impact.
What I learned
Reading a page was the easy part. The real engineering work was managing change, uncertainty and long-running work across a very large dataset.
For Capwell, the useful product was not a one-off scraper. It was a repeatable research and verification pipeline that could process a very large dataset, compare the result with existing records and involve a person only when the system needed judgment.
Each award source moved through extraction, normalization, comparison, verification and a controlled update path.
Confident matches continued to the database; ambiguous and uncertain results were routed to human review.
A comparable refresh moved from roughly six months to approximately three weeks across a 500K+ record database.
If your team is maintaining a large research, catalog or operations database by hand, I can help turn that process into an AI-assisted workflow with structured extraction, verification and human review built in.