Case study / Enterprise pilot / Healthcare

U.S. health insurance and services organization

Unlocking handwritten clinical data for agentic AI testing

The organization turned restricted clinical documents containing physician handwriting into a reusable, de-identified corpus for retrieval and agent evaluation.

01 / CHALLENGE

Why the existing data could not move.

Authentic clinical documents captured the handwriting, abbreviations, corrections, mixed content and layout variation that the AI platform would encounter. They also contained PHI and could not be placed in non-production environments or ingested into the platform's vector store. Fully synthetic text did not preserve enough of the source-document complexity, while manual redaction was slow and difficult to validate.

02 / APPROACH

What changed with GritWorks.

GritRedact processes the documents inside the organization's controlled environment, identifies PHI in typed fields, handwritten notes, annotations and other unstructured regions, and removes sensitive values while preserving approved clinical context, handwriting characteristics and document structure. Validated derivatives can then enter the controlled vector store and agent test workflow.

Results

What the organization can now do.

The outcome is not simply a redacted or synthetic document. It is a repeatable workflow that gives technical teams useful evidence without distributing the original sensitive record.

01

Converted sensitive clinical documents into a reusable, reviewed testing corpus.

02

Preserved physician handwriting, abbreviations, annotations and mixed-content layouts.

03

Removed member-identifying information before AI development and testing.

04

Enabled realistic testing of ingestion, extraction, chunking, embedding, indexing and retrieval.

05

Evaluated grounding, citation, abstention and escalation against known source evidence.

06

Reused the same validated corpus across model, prompt, embedding and vector-index updates.

Operational workflow

How the controlled data moves through the system.

  1. 01

    Submit a representative member-support question to the AI agent.

  2. 02

    Search the controlled vector store for relevant clinical documentation.

  3. 03

    Retrieve the most relevant document chunks.

  4. 04

    Interpret the approved typed and handwritten content.

  5. 05

    Summarize the evidence for the member advocate.

  6. 06

    Apply compliance guidance and human-escalation rules.

  7. 07

    Compare the response with the expected document, facts and outcome.

Next steps

Where the workflow expands next.

  • Expand to additional clinical-document types and handwriting styles.
  • Build larger golden datasets with expected sources and answers.
  • Measure retrieval accuracy across chunking, embedding and indexing strategies.
  • Automate regression testing when models, prompts or retrieval logic change.

Start with the blocked workflow

Make sensitive content usable—on terms your enterprise can defend.

Show us the source content, workflow and downstream users. We’ll identify where redaction, synthetic data or a different control is actually appropriate.

Request a working session