<!-- generated by scripts/generate-agent-docs.ts -->

> **Document Understanding Platform**
> Production HIPAA-compliant document extraction platform with multi-provider LLM integration, human-in-the-loop workflows, and human-feedback capture with reward-shaped labels for offline fine-tuning. Shipped at HCA Healthcare, the largest US hospital system.
>
> Source: https://jasonstiltner.com/projects/document-understanding/

---

*Part of my research on [coordination without collapse](https://jasonstiltner.com/corpus/): how do multiple AI providers coordinate on extraction tasks without losing the strengths of each?*

# Document Understanding Platform

Production document intelligence: HIPAA-compliant extraction with multi-provider LLM integration and human-in-the-loop learning

Updated Sep 9, 2026

HIPAA Compliant · Multi-Provider LLM · Human-Feedback Capture

[View on GitHub](https://github.com/jstiltner/document-understanding)

The repo is an open reference implementation of the architecture — the routing, confidence scoring, and feedback-capture code. It is not HCA Healthcare's production deployment and contains no PHI or protected code.

## The Challenge

Healthcare organizations process thousands of insurance authorization and denial documents daily—faxed PDFs, scanned forms, multi-page documents with varying quality. Manual extraction is slow, error-prone, and doesn't scale.

The challenge isn't just OCR—it's understanding document structure, extracting specific fields with high confidence, validating against business rules, and maintaining HIPAA compliance throughout. When extraction confidence is low, humans need to review and correct, and those corrections should improve future performance.

The architecture had to survive three constraints simultaneously, not sequentially: HIPAA-grade audit and access control, a volume and quality distribution shaped by HCA Healthcare's scale — scanners and fax gateways across the country's largest hospital system — and a reliability bar where a silent extraction error is worse than a slow one. Multi-provider routing with confidence-based escalation exists because no single model cleared that bar alone.

## System Architecture

[Diagram: Enterprise document understanding pipeline diagram: seven sequential stages — Document Upload, Quality Assessment, OCR Processing, LLM Extraction, Business Rules validation, Human Review, and Structured Output — plus a dashed reinforcement learning feedback loop returning human review corrections to the LLM extraction stage.]

## The Solution

Enterprise document processing platform with a complete extraction pipeline:

#### Multi-Provider LLM Integration

Seamless switching between Anthropic Claude, OpenAI GPT-4, and Azure OpenAI. Provider abstraction enables cost optimization, fallback handling, and A/B testing different models on the same documents.

#### Dual OCR Engine Support

Tesseract and EasyOCR with automatic quality assessment. Per-word confidence scoring enables intelligent routing—high-confidence extractions auto-complete, low-confidence documents route to human review.

#### Human-in-the-Loop Learning

Every human correction is captured with context as a reward-shaped label for offline fine-tuning. Reward shaping distinguishes between corrections, additions, and confirmations. Model performance tracked per field with precision/recall/F1 metrics.

#### Business Rules Engine

Configurable validation rules with severity levels. Cross-field validation catches logical inconsistencies (e.g., denial reason + authorization number). Custom Python expressions for complex business logic.

## Technical Implementation

### Multi-Provider LLM Service

The LLM service abstracts provider differences, enabling seamless switching and fallback. Field definitions are database-driven, allowing runtime configuration without code changes.

LLM Service

Multi-provider extraction with configurable field definitions

36 linesPythonFastAPI + Anthropic/OpenAI

View code

36 lines · Python · FastAPI + Anthropic/OpenAI

Copy

```python
class LLMService:
    """Multi-provider LLM extraction with graceful fallback."""

    def __init__(self, db: Session = None):
        # Initialize clients based on available API keys
        self.anthropic_client = Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
        self.openai_client = openai.OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
        self.azure_openai_service = AzureOpenAIService()

        self.default_provider = os.getenv("DEFAULT_LLM_PROVIDER", "anthropic")
        self.model_version = f"{self.default_provider}_{self.default_model}_v1.0"

    def extract_fields(self, ocr_text: str, provider: str = None) -> Dict[str, Any]:
        """Extract fields with configurable field definitions from database."""

        # Get field definitions (required vs optional)
        required_fields = self.field_service.get_required_fields()
        optional_fields = self.field_service.get_optional_fields()

        # Route to appropriate provider
        if provider == "azure_openai":
            return self.azure_openai_service.extract_fields(ocr_text, all_fields)
        elif provider == "anthropic":
            result = self._extract_with_anthropic(prompt, model)
        elif provider == "openai":
            result = self._extract_with_openai(prompt, model)

        # Calculate confidence scores per field
        confidence_scores = self._calculate_confidence_scores(extracted_data)

        return {
            'extracted_fields': extracted_data,
            'confidence_scores': confidence_scores,
            'requires_review': self._requires_review(extracted_data, confidence_scores),
            'model_version': self.model_version  # For RL tracking
        }
```

### Document Quality Assessment

Before extraction, documents are assessed for quality using computer vision metrics. Poor quality documents are flagged early with specific recommendations for improvement.

Quality Assessment

Multi-metric analysis with OpenCV and PIL

34 linesPythonOpenCV + Tesseract

View code

34 lines · Python · OpenCV + Tesseract

Copy

```python
class DocumentQualityService:
    """Multi-metric document quality assessment."""

    def assess_document_quality(self, file_path: str) -> Dict[str, Any]:
        # Image quality metrics
        quality_metrics = self._assess_image_quality(image_path, image)

        # Text quality via OCR confidence
        text_metrics = self._assess_text_quality(image_path)

        # Weighted overall score
        overall_score = self._calculate_overall_score(quality_metrics, text_metrics)

        return {
            "image_dpi": quality_metrics.get("dpi", 0),
            "image_clarity_score": quality_metrics.get("clarity_score", 0.0),
            "text_density_score": text_metrics.get("text_density", 0.0),
            "overall_quality_score": overall_score,
            "quality_issues": self._identify_issues(quality_metrics, text_metrics),
            "recommendations": self._generate_recommendations(...)
        }

    def _assess_image_quality(self, image_path: str, pil_image: Image) -> Dict:
        # Laplacian variance for clarity
        clarity_score = cv2.Laplacian(gray, cv2.CV_64F).var()

        # Contrast and brightness analysis
        contrast = gray.std()
        brightness_score = 1.0 - abs(gray.mean() - 128) / 128.0

        # Noise estimation via median filter comparison
        noise_level = self._estimate_noise_level(gray)

        return {"clarity_score": clarity_score, "contrast": contrast, ...}
```

### Business Rules Validation

Extracted fields are validated against configurable business rules. Cross-field rules catch logical inconsistencies that single-field validation would miss.

Workflow Service

Business rule validation with cross-field logic

41 linesPythonSQLAlchemy + Custom Rules

View code

41 lines · Python · SQLAlchemy + Custom Rules

Copy

```python
class WorkflowService:
    """Business rule validation and intelligent assignment."""

    def validate_business_rules(self, document_id: int) -> Dict[str, Any]:
        """Validate against all active business rules."""

        active_rules = self.db.query(BusinessRule).filter(
            BusinessRule.is_active == True
        ).all()

        violations = []
        for rule in active_rules:
            violation = self._validate_single_rule(document, rule)
            if violation:
                # Store in database for audit trail
                rule_violation = BusinessRuleViolation(
                    document_id=document_id,
                    rule_id=rule.id,
                    violation_details=violation,
                    severity=rule.severity
                )
                self.db.add(rule_violation)

                if rule.severity == "error":
                    violations.append(violation)

        return {"has_violations": len(violations) > 0, "violations": violations}

    def _validate_cross_field_rule(self, document, rule, rule_def) -> Optional[Dict]:
        """Cross-field validation (e.g., denial + auth number conflict)."""

        if rule_def.get("logic") == "denial_no_auth_number":
            denial_reason = document.extracted_fields.get("denial_reason")
            auth_number = document.extracted_fields.get("authorization_number")

            if denial_reason and auth_number:
                return {
                    "rule_name": rule.name,
                    "issue": "Denied documents should not have authorization numbers",
                    "severity": rule.severity
                }
```

### Human-Feedback Capture with Reward Shaping

Human corrections are captured with full context and written as reward-shaped training rows for an offline fine-tuning pass — confirmations are positive, corrections mildly negative, missed fields strongly negative. There is no preference model and no online policy update; this is labeled data collection, not RLHF.

Feedback Capture Service

Human-feedback capture with reward-shaped labels

45 linesPythonSQLAlchemy + custom reward shaping

View code

45 lines · Python · SQLAlchemy + custom reward shaping

Copy

```python
class ReinforcementLearningService:
    """Human feedback capture for continuous model improvement."""

    def record_human_feedback(
        self,
        document_id: int,
        field_name: str,
        original_value: str,
        corrected_value: str,
        feedback_type: str,  # correction, addition, removal, confirmation
        model_version: str
    ) -> HumanFeedback:

        # Calculate reward based on feedback type
        reward_score = self._calculate_reward(
            feedback_type, original_value, corrected_value
        )

        feedback = HumanFeedback(
            document_id=document_id,
            field_name=field_name,
            original_value=original_value,
            corrected_value=corrected_value,
            feedback_type=feedback_type,
            reward_score=reward_score,
            model_version=model_version,
            ocr_context=ocr_context  # For training context
        )

        # Update model performance metrics
        self._update_model_performance(model_version, field_name, feedback_type)

        return feedback

    def _calculate_reward(self, feedback_type: str, original: str, corrected: str) -> float:
        """Reward-shaped labels for offline fine-tuning."""
        if feedback_type == "confirmation":
            return 1.0   # Model was correct
        elif feedback_type == "correction":
            return -0.5  # Partial penalty (had something, was wrong)
        elif feedback_type == "addition":
            return -0.8  # Missed field entirely
        elif feedback_type == "removal":
            return -0.3  # Hallucinated field
        return 0.0
```

## Mathematical Formulation

Key algorithms underlying the document understanding pipeline.

### Per-Field Confidence Scoring

**Base Confidence:**

where \= extracted value for field , \= character length.

**Format Validation Boost:**

+0.1 boost if value matches expected pattern (dates, emails, phone numbers).

**Overall Document Confidence:**

where \= required fields (80% weight), \= optional fields (20% weight).

### Reward-Shaped Feedback Labels

**Feedback Reward Function:**

**Rationale:**

-   **Missed fields (-0.8)** penalized most: harder to catch in review than wrong values
-   **Corrections (-0.5)**: model found right location, wrong value
-   **Hallucinations (-0.3)**: easier to catch, less dangerous than omissions
-   **Confirmations (+1.0)**: positive reinforcement for correct extractions

**Per-Field Performance Tracking:**

where \= precision, \= recall for field , computed from human feedback over sliding window.

### Document Quality Assessment

**Image Clarity (Laplacian Variance):**

Higher variance indicates sharper edges. Blurry images have low .

**Brightness Score:**

where \= mean grayscale intensity. Optimal at 128 (mid-gray), penalizes too dark/light.

**Weighted Quality Score:**

Documents with are flagged with specific improvement recommendations.

## HIPAA Compliance Framework

Healthcare document processing requires strict compliance. The platform implements comprehensive safeguards:

#### Administrative

-   4-tier role-based access control
-   JWT authentication with expiration
-   Complete audit logging
-   Session timeout management

#### Technical

-   PHI access tracking per request
-   Security headers (HSTS, CSP, etc.)
-   Data integrity verification
-   Automated retention policies

#### Audit Trail Example

Every PHI access is logged with user, session, IP, action, resource, and data hash for integrity verification:

HIPAAAuditLog(
    user\_id="reviewer\_123",
    session\_id="sess\_abc",
    action="READ",
    resource\_type="document",
    resource\_id="doc\_456",
    phi\_accessed=True,
    ip\_address="10.0.1.50",
    data\_hash="sha256:a1b2c3..."
)

## What's Configured

These are counts and settings read out of the repository, not measurements. The system has never been load-tested — see [Design Trade-offs & Limitations](#trade-offs) for what that rules out claiming.

3

LLM providers, each with a real API client

15

SQLAlchemy models

100

`MAX_BATCH_SIZE` — documents accepted per upload

4

Default Celery `--concurrency` in docker-compose

### Technology Stack

#### Backend

-   FastAPI (async Python)
-   PostgreSQL + SQLAlchemy
-   Redis + Celery (queue)
-   Tesseract + EasyOCR

#### Frontend

-   React + TypeScript
-   Vite (build tooling)
-   Bootstrap components
-   Real-time updates

#### LLM Providers

-   Anthropic Claude (3 models)
-   OpenAI GPT-4/3.5
-   Azure OpenAI Service

#### Infrastructure

-   Docker + Docker Compose
-   Kubernetes-ready
-   Prometheus metrics
-   Swagger/OpenAPI docs

## Technical Deep Dive

Key architectural decisions and trade-offs—the questions a senior engineer would ask in a technical interview.

1 · How does the LLM provider abstraction handle different response formats?

Expand

Each provider has a dedicated extraction method (`_extract_with_anthropic`, `_extract_with_openai`) that normalizes responses to a common format. The key insight is that all providers return JSON, but with different wrapper structures:

-   **Anthropic:** `response.content[0].text`
-   **OpenAI:** `response.choices[0].message.content`
-   **Azure:** Same as OpenAI but with different auth headers

The `_parse_extraction_result()` method handles JSON cleanup (removing markdown fences) and field name normalization (mapping display names to internal names).

**Fallback logic:** If the primary provider fails, the system checks `get_available_providers()` and routes to the next available. Provider availability is determined by API key presence at initialization.

2 · How is per-field confidence calculated? What thresholds trigger review?

Expand

Confidence scoring uses a heuristic approach combining field completeness and format validation:

-   **Base confidence:** 0.8 for any non-empty extraction
-   **Format validation boost:** +0.1 if field matches expected pattern (dates, emails, phones)
-   **Length penalty:** Drops to 0.5 if value is <2 characters

**Overall confidence** is a weighted average: 80% from required fields, 20% from optional fields. This ensures missing required fields heavily impact the score.

Review Triggers (configurable via env vars):

-   `MIN_CONFIDENCE_THRESHOLD=0.7` — Overall confidence below this triggers review
-   `REQUIRED_FIELDS_THRESHOLD=0.8` — Any required field below this triggers review
-   Missing required fields — Always triggers review regardless of confidence

3 · How is human feedback captured, and how is it used for improvement?

Expand

The system captures human feedback with full context for a future offline fine-tuning pass — not a live training loop:

Reward Shaping:

-   **Confirmation (+1.0):** Human verified model was correct
-   **Correction (-0.5):** Model extracted something, but wrong value
-   **Addition (-0.8):** Model missed field entirely (false negative)
-   **Removal (-0.3):** Model hallucinated a field (false positive)

**Why these specific values?** Missing fields (-0.8) is penalized more than wrong values (-0.5) because a wrong extraction at least indicates the model found the right location. Hallucinations (-0.3) are penalized less because they're easier to catch in review than missed fields.

**Data captured for training:** Each feedback record includes the original OCR context, model version, field name, original/corrected values, and reviewer ID. This enables:

-   Fine-tuning on correction pairs (original → corrected)
-   Prompt engineering improvements based on common error patterns
-   Per-field performance tracking to identify weak spots

4 · What specific HIPAA safeguards are implemented? How is PHI tracked?

Expand

HIPAA compliance is implemented via middleware that intercepts every request:

-   **PHI endpoint detection:** Routes like `/documents`, `/upload`, `/review` are flagged as PHI-accessing
-   **Audit logging:** Every PHI access creates a `HIPAAAuditLog` record with user, session, IP, action (CRUD), resource type/ID, and data hash
-   **Data integrity:** SHA-256 hash of access metadata enables tamper detection

Security Headers (added to every response):

-   `X-Frame-Options: DENY` — Prevents clickjacking
-   `Strict-Transport-Security: max-age=31536000` — Forces HTTPS
-   `Content-Security-Policy: default-src 'self'` — Prevents XSS
-   `X-Content-Type-Options: nosniff` — Prevents MIME sniffing

**Role-based access:** 4-tier system (Admin, Supervisor, Reviewer, Viewer) with permission checks before PHI access. Supervisors can access quality checks; Reviewers can only access assigned documents.

5 · What quality metrics are used? How does quality affect routing?

Expand

Document quality assessment uses computer vision metrics before OCR:

Image Quality Metrics:

-   **DPI:** Extracted from image metadata, minimum 200 recommended
-   **Clarity:** Laplacian variance (cv2.Laplacian) — higher = sharper
-   **Contrast:** Standard deviation of grayscale values
-   **Brightness:** Distance from optimal mean (128) — penalizes too dark/light
-   **Noise:** Difference between original and median-filtered image

Text Quality Metrics (via OCR):

-   **Text density:** Ratio of detected text blocks to total blocks
-   **Average confidence:** Mean OCR confidence across all words
-   **Readable ratio:** Percentage of words with confidence >60%

**Weighted overall score:** DPI (20%), Clarity (25%), Contrast (15%), Brightness (10%), Noise (10%), Text confidence (15%), Text density (5%). Documents below `MIN_QUALITY_THRESHOLD=0.5` are flagged with specific recommendations (e.g., "Scan at higher resolution", "Ensure document is in focus").

### Design Trade-offs & Limitations

Retracted · Sep 2026 history

This page previously claimed “500+ documents/hour (4 workers).” That number was never measured — there is no load test, benchmark script or results artifact in the repository. It also contradicts the repo's own README, which puts single-document processing at 30–60 seconds; four workers at the optimistic end of that range caps out near 480/hour, and near 240/hour at the pessimistic end. The “100+ concurrent” figure was likewise the configured `MAX_BATCH_SIZE` rather than an observed concurrency.

**No automated tests:** the repository has no test suite and no CI workflow — zero test files, zero `def test_`. Everything described here was verified by reading the code, not by running it. For a system whose selling point is auditability in a regulated setting, that is the most significant gap on this page.

**Accuracy figures are deliberately absent:** the README quotes “95%+ OCR accuracy,” “85%+ field extraction” and a “<5% false positive rate.” None of those are reproduced on this page, because nothing in the repository measures them. They should not be imported here later.

**Heuristic confidence vs. model-based:** Current confidence scoring uses rule-based heuristics rather than a trained confidence model. This was a deliberate choice for interpretability and ease of tuning, but a learned confidence estimator could be more accurate.

**No online learning loop:** Feedback is captured but not used for real-time model updates. The system generates reward-shaped training data for offline fine-tuning rather than online learning — so this is feedback capture, not RLHF. Appropriate for healthcare, where model changes require validation before they ship.

**OCR-first architecture:** The system OCRs entire documents before LLM extraction. An alternative would be vision-language models (GPT-4V, Claude 3) that process images directly. The OCR-first approach was chosen for cost efficiency and to leverage existing Tesseract infrastructure.

## Why It Matters

This project demonstrates several capabilities relevant to production ML teams:

-   **Production ML Engineering:** Not a prototype—full-stack application with database models, async processing, monitoring, and deployment configuration.
-   **LLM Application Architecture:** Multi-provider abstraction, prompt engineering for structured extraction, confidence scoring, and graceful degradation.
-   **Human-in-the-Loop ML:** Systematic capture of human feedback with reward-shaped labels — the input to an offline fine-tuning pass, and the dataset an RLHF loop would later require.
-   **Healthcare Domain Expertise:** HIPAA compliance isn't an afterthought—it's built into the architecture with comprehensive audit logging and access controls.

#### Portfolio Context

This project demonstrates healthcare ML expertise in document processing and extraction. It shows depth in the healthcare AI stack—from unstructured document ingestion to structured data extraction with HIPAA compliance.

[See Research Project: Collaborative Nested Learning](https://jasonstiltner.com/projects/collaborative-nested-learning/)
