Feel free to get in touch with us for more information about our products and services.

Every year, pension funds and insurers pay out benefits to people who are no longer alive to receive them. The infrastructure meant to catch a death hasn’t caught up yet, so checks keep getting cut and policies stay active while the record goes unverified.
The Social Security Death Master File (DMF) is the default source that most financial institutions usually rely on. However, it lags the actual date of death by 4 to 6 months. In that window, checks can keep getting cut, policies stay active, and beneficiary records go unverified.
To close that gap, death audit and locator service providers work with pension funds, insurers, banks, and healthcare providers.
How? With a simple insight:
Obituaries are often published within days of a death, long before any government database catches up. If you could reliably turn obituary listings into structured, searchable data, you’d have a fraud and compliance signal faster than almost anything else available.
But there’s a catch, obituaries are not built to be structured data. They’re prose, written by grieving families, published on funeral home and memorial sites with no consistent format.
Getting a name, a date of birth, or a service location out of that text reliably, and at scale, meant solving a very different kind of extraction problem than most scraping projects Grepsr faces.
This is a story about a client that is a death audit and locator service provider. They work with pension funds, insurers, banks, and healthcare providers to catch deaths faster than traditional record sources.
Their business depends on closing exactly the kind of gap described above: the earlier a death is confirmed, the sooner a benefit payment stops, a policy gets flagged, or a beneficiary record gets updated.
They came to Grepsr with a specific, two-source obituary data project:
Each obituary had to be turned into structured and usable fields:
Full name; date of birth, date of death, and age; the obituary text itself; the source URL; funeral home name and address; funeral service date and location, and more.
Usually, the data extraction projects Grepsr gets are structured, with prices in price fields, product names in title tags, and everything else following a predictable layout. Obituaries, on the other hand, don’t work that way.
They’re a short piece of prose written by a grieving family member or funeral home staffer. The facts a fraud-detection system actually needs, like date of birth, date of death, and age, are buried inside sentences rather than sitting in labeled fields.
That created a few distinct extraction problems.
Full names on the page came as a single string, sometimes with a nickname in quotes, a maiden name in parentheses, or a suffix. Splitting that reliably into first, middle, and last name meant handling different formats for each individual obituary.
A date of birth or death might appear as “born on March 4, 1952” in one obituary and “entered eternal rest on 3/4/2021 at the age of 68” in another. There was no consistent structure to hook into, so DOB, DOD, and age all had to be parsed out of free text with enough reliability to trust downstream.
A single obituary could reference several distinct events: a visitation, a funeral mass, a graveside burial, each with its own date and location. Some records in the dataset carried as many as six separate events. Each one needed to be captured as its own structured entry rather than flattened into a single field.
Obituary sites present this information differently, so the extraction logic had to normalize funeral home name, address, and “published by” text consistently across two sites that don’t share a layout.
What especially made it hard was doing it at scale, across two different sites, on a recurring daily basis, without the parsing breaking every time an obituary was worded a little differently than the last one.
Grepsr built two distinct obituary data collection strategies matched to two distinct goals, rather than treating both sites as a single uniform scrape.
With this ongoing project, now our client, the death audit service provider can
The most valuable data isn’t always sitting in a structured database somewhere, waiting to be queried. Sometimes it’s published every day, in plain text, by people who have no idea they’re creating a signal anyone else needs.
Obituaries were never designed to be a fraud detection tool. But once that free text gets parsed, structured, and delivered reliably, it becomes exactly that: a faster, more human source of truth than the records built to replace it.
This project is also a reminder that web scraping, done well, isn’t just pulling records off a page. It’s deriving meaning from what’s already public. Anyone can grab the text of an obituary.
Turning a grieving family’s words into a clean, structured, trustworthy dataset, at scale, every single day, is a different problem entirely. That’s the part that actually matters.
If your business depends on knowing something before everyone else does, the data you need might already be public. It just hasn’t been scraped and structured yet.
Talk to Grepsr about a tailored scraping solution built to surface the signals your business can’t afford to miss. \