Own project, rebuilt on made-up data

Messy listings in, clean company records out.

My first job was collecting company details from public business listings by hand and typing them into records. Later I built the parser that did it: it reads each listing however it is written, pulls out the fields, normalises the phone numbers and websites, and merges the same company listed twice. Here it is on invented listings.

The companies below do not exist. The listing formats are the kind you meet in real directories.

About 40 listings in three directory formats, with a few duplicates and a few holes.
Show the raw listings as they come

What the real version did.

The same, at the scale of thousands of listings a week, with one parsing profile per directory and a review queue for the rows it was not sure about. The research team went from typing to checking, which is a different job and a much faster one.

If someone on your team copies company details from the web into a spreadsheet, the first project is a parser for your main source, from €400. Show me the source ›