pipeline/spark_ingestion.py -> run_ingestion() Scans the input folder, distributes PDFs across Spark partitions, extracts text, chunks it, embeds it, and uploads documents to the selected Azure AI Search index. Text extraction pipeline/spark_ingestion.py -> extract_text() Uses PyMuPDF / pymupdf4llm for text PDFs. If is_scanned() finds very little embedded text, it falls back to Azure Document Intelligence OCR. Chunking pipeline/spark_ingestion.py -> chunk_text() Uses token chunking by default; semantic chunking is available through SemanticChunker. Each chunk keeps metadata such as paper_id, doc_type, regulation, page, and chunk_id. Framework KG build pipeline/kg_builder.py -> ComplianceKGBuilder.build_from_pdfs() Reads framework PDFs from data/regulations, splits them into article/section chunks, runs LLMGraphTransformer, and writes graph documents to Neo4j. KG node/edge schema pipeline/kg_builder.py -> ALLOWED_NODES, ALLOWED_RELATIONSHIPS Nodes include Regulation, Article, Obligation, Right, Entity, Concept, Penalty, and Timeframe. Edges include REQUIRES, GRANTS, REFERENCES, MAPS_TO, PART_OF, DEFINES, APPLIES_TO, IMPOSES, CONFLICTS_WITH, and STRICTER_THAN. Value conflict extraction pipeline/kg_builder.py -> _extract_value_conflicts() Runs a dedicated LLM pass over framework text and writes concrete CONFLICTS_WITH / STRICTER_THAN edges using Cypher. Stored properties include concept, value_a, value_b, unit, description, and source_quote. Policy selection agent/compliance_nodes.py -> doc_resolver_node() Uses the frontend dropdown value first. If absent, it tries filename/query matching; if multiple policies are possible, it returns a clarification response. Framework detection agent/compliance_nodes.py -> jurisdiction_detector_node() Detects the applicable regulatory frameworks from the question and available context. These names are then used to scope obligations, conflicts, and scoring. Evidence retrieval agent/compliance_nodes.py -> kg_retriever_node() and pipeline/compliance_retriever.py -> _azure_search() Retrieves policy chunks from Azure AI Search with a filter like doc_type eq 'policy' and paper_id eq '<selected.pdf>'. It also retrieves framework chunks from the framework index. KG context retrieval pipeline/compliance_retriever.py -> multi_hop() Extracts important terms from the retrieved text, then asks Neo4j for nearby connected framework facts. In plain English: if the query mentions breach notification, the graph helps pull related duties, timeframes, rights, penalties, and framework sections connected to that concept. Obligation structuring agent/compliance_nodes.py -> _triples_to_obligations() Converts Neo4j relationship results into simple obligation records with framework, obligation type, source, and text fields. Keyword rules classify obligations such as breach notification, data subject rights, access control, encryption, retention, consent, and lawful basis. Gap analysis agent/compliance_nodes.py -> gap_analyzer_node() and _analyze_gaps_with_llm() Sends the selected policy text plus structured obligations to the LLM in batches. The LLM returns JSON gaps with severity, framework, obligation ID, theme, evidence, and missing/weak policy language. Conflict filtering agent/compliance_nodes.py -> conflict_detector_node() Reads CONFLICTS_WITH and STRICTER_THAN edges from Neo4j and keeps only conflicts relevant to the detected frameworks for that audit. Priority scoring agent/compliance_nodes.py -> risk_scorer_node() Converts gaps into a simple prioritization score for the UI/report. Severity weights are critical=10, high=7.5, medium=5, low=2.5, info=1. Framework weights are GDPR=0.35, HIPAA=0.30, CCPA=0.20, NIST=0.15. The displayed compliance score is calculated by risk_to_compliance() = 100 - risk * 10, clamped to 0-100. This is not a legal certification score; it is a way to rank and summarize findings. Remediation agent/compliance_nodes.py -> remediation_node() and _generate_remediations() Groups gaps by closed-vocabulary themes and asks the LLM for one concrete checklist action per theme. The final report renders these as remediation checklist items. Runtime telemetry app/core/telemetry.py and agent/compliance_nodes.py Emits visible console logs and Application Insights traces for query_received, node_jurisdiction_detector, node_gap_analyzer, node_risk_scorer, and query_completed.pipeline/spark_ingestion.py file handles extraction, scanned-PDF OCR fallback, chunking, embeddings, and Azure AI Search upload..env file:frontend/.env:VITE_API_URL to the public backend URL and restart npm run dev.GET /api/v1/health Liveness check. GET /api/v1/health/ready Readiness check. POST /api/v1/query Run an audit query. POST /api/v1/query/stream Run an audit with streaming progress. POST /api/v1/ingest Upload policy PDFs. GET /api/v1/policies List indexed policy documents. GET /api/v1/history Read MLflow-backed audit history when configured. GET /api/v1/conflicts Read framework conflicts from Neo4j. GET /api/v1/regulations List indexed regulatory frameworks. GET /api/v1/regulations/chunks Browse regulatory-framework chunks.mlops/compliance_tracker.py.Posted Sep 22, 2026
Developed a compliance tool for auditing privacy policies using AI and other technologies.