Duplicate removal

Remove duplicate pages from medical records

Duplicates are the worst part of reading a medical record set. RecordFlow finds them and drops them: the page faxed three times, the report both the hospital and the clinic sent you, the whole file someone attached twice. What comes back is one PDF in date-of-service order.

What counts as a duplicate page?

A duplicate is the same page of the same record arriving more than once. In a personal-injury file that happens constantly: a provider faxes the same visit note twice, the hospital and the imaging center both send the same radiology report, a paralegal attaches the same PDF to two emails, and a 900-page production turns out to hold 200 pages you have already read. RecordFlow removes all four kinds.

How does RecordFlow find duplicate pages?

Two passes. First, any uploaded file that is byte-for-byte identical to another is dropped whole, so the same PDF sent twice costs you nothing. Then every remaining page is read with OCR, and pages sharing a date of service are compared on their text: exact matches, matches that differ only in OCR noise, and pages that are close but not identical. Fax headers, portal timestamps and printed-on footers are stripped before the comparison, because those are the parts a second copy always changes.

Does it catch duplicates across different providers and different files?

Yes. Pages are compared across every file in the case, not within each file separately. A page in the hospital's production and the same page in the orthopedist's production are compared against each other, which is the case that matters most, because that is the duplicate a person reading one file at a time can never see.

What happens when it is not sure two pages are the same?

It keeps both. Pages that are similar but not clearly identical go to a second, separate AI check that looks at the full surrounding records rather than the two pages alone. If that check fails, times out, or returns anything its validators do not accept, the result is keep everything. A page is only ever dropped by a check that succeeded and passed every validator. Losing a real page is a far worse outcome than leaving a duplicate in, and the engine is built to fail in that direction.

Is this HIPAA-compliant?

Records are processed only on cloud infrastructure covered by signed Business Associate Agreements (AWS for hosting and OCR, Google Cloud for Vertex AI). Files are encrypted in transit and at rest, uploaded records and the organized output are permanently deleted 30 days after processing, and records are never used to train any AI model. A BAA for your firm is available on request.

What does duplicate removal not do?

It compares the text on the page, not the image, so a scan too degraded for OCR to read may not match its own twin. It is page-level deduplication of one patient's records, not the master-patient-index matching that hospital systems mean by duplicate records. Nobody at RecordFlow reads your file to arbitrate a close call, and automated output can contain errors, so review the record before relying on it.

What do I get back?

One PDF containing every remaining medical page in chronological order by date of service, across every provider, with duplicate and non-medical pages already gone. Pricing is per case by page count, with no subscription, and the price is confirmed with you before anything runs. Your first case is free, so you can see what comes back on your own records first.

The rest of the process is on How it works, the security model in full is on Security, and the per-case rates are on Pricing. For a BAA or a question about a specific file, email support@s2reason.com.

Send us a case →