Evidence-Grounded AI Research Audit + by Phillip WellsEvidence-Grounded AI Research Audit + by Phillip Wells

Evidence-Grounded AI Research Audit +

Phillip Wells

Phillip Wells

Evidence-Grounded AI Research Audit + VERA Company Dossier — AI Evaluation & Company Research You Can Trust
I test whether your AI actually did the research — then I build the reusable skill that fixes the errors it keeps making.
THE PROBLEM
The most dangerous AI failure isn't an obvious error. It's a confident, well-written company or market profile that was never grounded in a source — wrong products, missing competitors, suppliers listed with no context, contradictions smoothed over. Teams and investors make decisions on that.
WHAT I BUILT — THE AUDIT SYSTEM
An evidence-grounded audit system designed to detect AI that sounds researched without actually researching:
Provenance controls — key claims must trace to an actual source
A decision-critical unknown gate — if something that would change the decision is unknown, the system must say so instead of filling the gap
Functional-role classification and real substitutability checks
Disconfirming search — actively looking for evidence against the leading conclusion
Prior-correction Regression Tests — once an error is caught, the system is retested so it doesn't quietly come back sier (v1.0)** To fix a good bit of those errors in company research specifically, I turned the audit principles into a production skill in my Drive Skills library — vera-company-dossier.
~Its stated job is to produce a company overview you can trust as if you'd researched it from first principles yourself.
~It is aimed squarely at the recurring failure it was built to solve: a model finds a company's most familiar historical identity, reads a handful of summaries, and mistakes that partial picture for the company.
HOW THE SKILL ACTUALLY WORKS (verified from the skill itself):
13 Mandatory Research Passes — entity/freshness lock, identity reconstruction, strategic history, product & technology census, revenue/economic engine, customer/partner/supplier graph, architecture/dependency role, competitive landscape, macro/policy/supply-chain exposure, management & capital allocation, valuation, catalysts/risks/falsifiers, and a mandatory blind-spot expansion sweep
Legacy Label vs. Current Operating Identity — every dossier must explicitly compare what the market calls the company with what it actually is today, and reject the label if current evidence can't support it
A 3-tier Source Ladder — Tier 1 primary evidence (SEC filings, earnings, investor-day materials, official product docs, counterparty confirmation, government records) outranks Tier 2 independent reporting, and Tier 3 (analyst posts, blogs, social, search snippets, prior model answers) is discovery leads only
An 8-state Evidence Vocabulary for Every Material Claim — Verified Current, Counterparty Verified, Reported, Announced / Not Yet Proven, Historical, Inferred, Disputed, or Unknown. An announced roadmap item, pilot, MOU or design win is never upgraded to shipped revenue
Hard Anti-failure Rules — don't confuse technical credibility with commercial maturity, backlog/TAM with realized profit, or a partnership announcement with material revenue; don't let a hype narrative erase the current revenue engine, or the revenue engine hide a new platform
A backlog-quality Audit and Capital-structure/dilution Normalisation — headline backlog is decomposed into firm vs. conditional vs. non-binding, and valuation is reconciled to the real ownership base (warrants, convertibles, SBC, post-offering cash). Both were added permanently after the skill's LEU / Centrus Energy acceptance test on 2026-10-04 exposed them
A Blind-spot Sweep with a Discovery-saturation Rule — second-order searches built from products, subsidiaries, partners, acquisitions and standards must run until two consecutive passes find nothing new and material
A Contradiction Gate and a 15-domain Completeness Matrix — a dossier may only report DOSSIER COMPLETE when no decisive domain is Unknown, no material contradiction is hidden, the latest filings and earnings were checked, relationships were counterparty-checked where possible, and the blind-spot sweep reached saturation. Otherwise it must output DOSSIER INCOMPLETE with the exact gaps. No amount of eloquence overrides the gate
Four Working Modes — Full Dossier, Blind-Spot Audit, Dossier Refresh, and Pre-Decision Research
Calibration / Regression Cases built in — QCOM (legacy label hiding a compute continuum), Lenovo (PC label vs. enterprise AI/edge systems), XNDU (technically important, commercially early), and LEU (single-theme label vs. a multi-path bottleneck company)
The dossier researches and verifies the company; it does not authorize a trade or a decision. Research authority and decision authority stay separate, and a completed dossier hands its evidence to whatever decision process comes next.
WHAT I CAN DO FOR A CLIENT
Run a structured reliability audit of your AI assistant, research tool, or agent workflow
Build a company dossier / deep-dive research package — for investment diligence, partnership evaluation, competitive research, or pre-decision company understanding
Build a reusable dossier/research skill for your own team, with the completeness gates baked in
Build an evaluation / test harness for your AI: test cases, failure taxonomy, scoring, and regression checks
Diagnose why your AI is hallucinating or over-claiming — which layer is actually failing (sources, reasoning, instructions, retrieval, or interface)
Like this project

Posted Oct 5, 2026

Evidence-Grounded AI Research Audit + VERA Company Dossier — AI Evaluation & Company Research You Can Trust I test whether your AI actually did the research ...