Dealership data enrichment: one confident answer per rooftop, or none.
For an automotive dealer network (name withheld), Aldenebai built a dealership data enrichment pipeline: software that works out, rooftop by rooftop, which service-scheduler vendor each dealership runs and writes the answer into HubSpot only when sure. It asks OEM APIs first, then vendor subdomain enumeration and reverse lookups, and opens the dealer's own website only as a last resort. Only HIGH-confidence results reach HubSpot's dms field, and nothing was mass-written until the pipeline cleared 95% precision on a hand-labelled test set.
Dealer data at the network before the pipeline
One fact was needed for every rooftop in HubSpot: which service scheduler takes the store's bookings. Xtime, Dealer-FX, CDK, TotalCustomerConnect, or one of the others. The manual way is a researcher opening each dealer's website, hunting for the booking widget, and typing a vendor name into the CRM. It is slow, inconsistent, and built on the least reliable source there is: old widgets, cached pages, templates reused across a group. The existing dms field mixed hand-typed answers with values copied from the group. Some of it was right; nobody could say which part.
How the enrichment pipeline resolves each rooftop’s scheduler
Aldenebai built a service scheduler vendor detection pipeline that treats its sources in order of authority. OEM APIs are asked first; an answer there closes the search for that dealer. If the OEM is silent, it enumerates vendor subdomains belonging to the dealer, then runs reverse lookups on what it finds. Only when all of that comes up empty does it open the dealer's own site, and what it sees there is scored rather than believed. HIGH-confidence results are written to HubSpot's dms field; anything lower stays a lead or is flagged for a person with the evidence attached. Resolution is per dealer only: a rooftop never inherits a vendor from its group. Before the first mass write it had to reach at least 95% precision on a hand-labelled test set, and it removed the false positives the manual process had left in the prior data.
What the research team does now
The researcher's job changed shape. Instead of opening dealer sites one by one, they review the flagged rooftops, each arriving with the OEM response, the subdomains found, and the lookups already run, and confirm or reject with the facts in front of them. The dms field is now something the CRM can build on: HubSpot and CRM automation runs deduplication, routing, and segmentation off a value that was earned, not guessed.
Results: high-confidence writes only, false positives cleaned out first
The 95% figure is a hard gate, not an estimate.
Book the Automation Audit
The receipt
Illustration of the system's output, not a production screen
What people ask about this dealership data enrichment case study
How does service scheduler vendor detection work when the dealer website is misleading?
By not starting there. The pipeline asks OEM APIs first, then enumerates vendor subdomains and runs reverse lookups, and opens the dealer site only when those sources are silent. Whatever the site shows is scored, not trusted. A stale booking widget never earns a HIGH.
What does HubSpot enrichment with confidence gates mean in practice?
Every answer carries a confidence level, and only HIGH is ever written to the dms field. Anything below that stays a lead or is flagged for a human with the evidence attached. On top of that per-record gate sits a batch gate: at least 95% precision on a hand-labelled test set before any mass write was allowed.
Why is the answer resolved per dealer instead of per dealer group?
Because inheriting from the group is how false positives get in. Each rooftop is resolved on its own evidence and gets its own result or none at all. A group-level assumption looks safe until it is wrong, and those are exactly the records nobody goes back to fix.
What happened to the data that was already in HubSpot?
It was cleaned before the pipeline wrote anything new. Prior false positives were removed from the dms field, so old guesses do not sit next to new evidence. Every value left can be traced to the source that produced it.
Related pages
Your dealer list is the next case.
The free Automation Audit maps how vendor data gets into your CRM and where a confidence-gated pipeline pays for itself first — 30 minutes, written Automation Map within 48 hours.