Slides: AI-Driven Multi-Document Correlation for Financial Compliance - Varsha Shah, Independent
Source Video
AI-Driven Multi-Document Correlation for Financial Compliance - Varsha Shah, Independent
Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.
Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.
Extracted Slides

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
title_cardconfidence0.99 - Text source: agent_vision.
Slide text:
AI-Driven Multi-Document Correlation for Enterprise Financial Compliance and Fraud Detection
A framework for cross-document fraud detection through relational intelligence, evaluated across 3 million anonymized records and four jurisdictions.
By Varsha Shah, Enterprise Technical Architect, USA

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: agent_vision.
Slide text:
The Compliance Gap No One Is Closing
Multi-Jurisdictional Complexity
Growing Data Volumes
Sophisticated Fraud Patterns
Rule-based and NLP-augmented systems operating at the document level are structurally incapable of detecting these cross-document anomalies.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: agent_vision.
Slide text:
Why Document-Level Analysis Falls Short
Traditional compliance tools evaluate records in isolation. Fraud that arises from discrepancies between payroll registers, vendor invoices, and tax filings remains invisible when each document passes its own internal validation.
The most costly fraud patterns are not found within a single document. They emerge in the space between documents.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: agent_vision.
Slide text:
Framework Architecture Overview
Entity Correlation
Risk Modeling
Normalization Layer
The three components operate in concert: entities are linked relationally, risk signals are aggregated and calibrated, and jurisdictional variance is normalized before scoring.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.96 - Text source: agent_vision.
- OCR decision: ready — Dense text slide with title, paragraph, and bullets; OCR is appropriate.
Slide text:
Graph-Based Entity Correlation Engine

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.96 - Text source: agent_vision.
- OCR decision: ready — Dense text slide with title, paragraph, and bullets; OCR is appropriate.
Slide text:
Adaptive Probabilistic Risk Model

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: agent_vision.
Slide text:
Evaluation Conditions
Dataset Scale
Approximately 3 million anonymized financial records
Jurisdictions
Four distinct regulatory environments evaluated in parallel
Time Horizon
Five years of historical data reflecting real-world enterprise conditions

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: agent_vision.
Slide text:
Detection Performance Results
~91% Precision
~87% Recall
~0.89 F1 Score
Performance was measured against a labeled ground truth derived from confirmed audit findings across all four jurisdictions and the full five-year evaluation window.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: none.
- OCR decision: ready — Dense slide with charts, captions, and small body text.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: none.
- OCR decision: ready — Dense multi-column comparison slide with small text.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: none.
- OCR decision: ready — Diagram slide with multiple labels and a paragraph of small text.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: none.
- OCR decision: ready — Four dense text panels with small body copy.

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
title_cardconfidence0.99 - Text source: agent_vision.
Slide text:
Thank you.
Varsha Shah — Enterprise Technical Architect
linkedin.com/in/varsha-shah-7b5111247
varsha.shah.tech@gamil.com
Classification audit: raw/sources/slide-ai-classification/slides/Iwe_RY-fYgI/audit.json
Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.