Master and Reference Data Management Service

Resolve fragmented identities into trusted, governed enterprise records

★★★★★4.9 out of 5 from 6,247 reviews

Dataconsultant helps organisations identify when records refer to the same customer, supplier, patient, citizen, employee or other entity. We assess source data, define deterministic and probabilistic matching, design survivorship and stewardship controls, support implementation, and establish measurable operations for trusted identity views across systems.

  • Explainable matching and confidence rules
  • Governed survivorship and exception handling
  • Privacy, security and consent considerations
  • Vendor-neutral architecture and implementation support

What is identity resolution?

Identity resolution is the controlled process of determining which records across one or more systems represent the same real-world entity, then linking or consolidating them according to approved matching, survivorship and governance rules. It is commonly sponsored by data, technology, operations, customer, risk or compliance leaders when duplicate and fragmented records undermine service, analytics, fraud controls, reporting or regulatory evidence. Typical outputs include a source assessment, matching strategy, rulebook, golden-record model, exception workflow, implementation design and KPI framework. Results depend on source quality, lawful data use, business ownership and sustained stewardship; the service does not replace legal advice or formal assurance.

Service offering

Assessment, solution design and operational enablement

The engagement can address a focused identity problem or establish an enterprise capability spanning data, technology, governance and operations.

01

Assess identity risk and readiness

Profile source systems, identifiers, duplicates, conflicts, missing values, consent constraints, lineage, existing rules and operational pain points. Inputs include samples, schemas, policies, issue logs and stakeholder knowledge. Outputs include quantified findings, candidate match keys, data-quality dependencies, risks and scope recommendations.

02

Design matching and golden records

Define entity boundaries, match strategies, thresholds, survivorship, source precedence, householding, relationship rules, manual review, audit evidence and target architecture. Client owners validate definitions and risk tolerances. Outputs include an approved rulebook, data model, workflow and implementation backlog.

03

Implement, validate and operate

Support configuration or engineering, test datasets, tuning, reconciliation, controls, deployment, stewardship training and service transition. Delivery may use an existing MDM platform, cloud service, data platform or custom components. Value is sustained through monitoring, accountable ownership and change control.

Key value propositions

Practical value from a reliable view of identity

More consistent customer and entity views

Link records across channels and systems using controlled evidence rather than unsupported assumptions, improving the foundation for service, reporting and analysis.

Lower duplicate-driven operational friction

Reduce repeated outreach, conflicting updates, manual reconciliation and fragmented case handling where identity data is a material cause.

Stronger governance and explainability

Document why identities were linked, which attributes survived, who approved exceptions and how rules changed over time.

Better risk and privacy control

Incorporate lawful-use constraints, sensitive-data handling, access controls, retention, consent and data-minimisation requirements into design decisions.

Problems addressed

Where identity fragmentation creates business and control risk

Identity issues rarely sit in one database. They affect workflows, decisions, customer experiences, controls and the credibility of enterprise reporting.

Duplicate records across systems

Impact: Teams contact the same person repeatedly, maintain conflicting attributes and spend time reconciling reports.

Response: Dataconsultant profiles duplicates, defines match evidence and establishes merge, link or exception decisions with provenance.

Inconsistent or changing identifiers

Impact: Email, phone, address, account and device identifiers may be missing, reused or changed, making exact joins unreliable.

Response: We combine deterministic, fuzzy and probabilistic techniques based on risk, available evidence and acceptable false-match tolerance.

Unclear survivorship and source authority

Impact: Conflicting names, addresses, classifications or statuses create uncertainty about which value should be trusted.

Response: We define field-level precedence, recency, validation, source reliability and stewardship rules rather than applying one blanket rule.

Unexplained automated links

Impact: Black-box decisions can create privacy, fairness, service and audit concerns, particularly in regulated or high-impact processes.

Response: We design confidence bands, reason codes, review thresholds, audit trails and quality checks suited to the consequence of an incorrect match.

Fragmented consent and preference records

Impact: Linking identities without considering purpose, consent or restrictions can propagate data into inappropriate channels.

Response: Privacy and consent requirements are mapped into identity models, permitted uses, access controls and downstream distribution rules.

Identity models that do not scale

Impact: Point solutions may fail when volumes, sources, jurisdictions or entity types increase.

Response: We assess batch and real-time needs, performance, model governance, monitoring, stewardship capacity and integration dependencies.

Clarify whether identity resolution is the right intervention

Review your entity types, source systems, operational consequences and risk tolerance with a specialist.

Request a Consultation
Suitability

Who identity resolution is for

The service can support startups, growing businesses, enterprises, public-sector bodies and regulated organisations where multiple records must be connected responsibly.

Good fit

  • Multiple systems hold overlapping records for the same entity.
  • Duplicate or fragmented identities affect service, analytics, fraud, compliance or reporting.
  • A merger, migration, CRM programme, MDM initiative or customer-data programme requires consolidation.
  • Existing matching rules are undocumented, inaccurate or difficult to govern.
  • Business owners can define entity meaning and acceptable error tolerances.
  • Data, privacy, security and operational stakeholders can participate.

May not be the right fit

  • A small data-cleaning exercise can resolve a one-off duplicate list.
  • A broader process, operating-model or platform transformation is the primary need.
  • A software product alone can satisfy a well-defined, low-risk requirement.
  • A permanent internal data stewardship or engineering hire is more appropriate.
  • A licensed legal opinion, statutory audit or specialist cybersecurity test is required.
  • Necessary data access, ownership decisions or lawful-use approvals are unavailable.
Common use cases

Identity resolution across different operating contexts

Omnichannel customer identity

Connect CRM, ecommerce, service, loyalty and marketing records while preserving consent and source provenance.

Scope: Customer matching and golden profile
KPIs: Duplicate rate, review rate, precision
Model: Design plus implementation
Dependency: Consent and identifier quality

Supplier and counterparty consolidation

Identify the same legal or trading entity across procurement, finance and risk systems to support spend visibility and due diligence.

Scope: Legal entity and hierarchy resolution
KPIs: Duplicate suppliers, unmatched spend
Model: Assessment and roadmap
Dependency: Registration identifiers

Patient or member matching

Improve record linkage across care, claims or membership systems with stricter thresholds, review controls and privacy safeguards.

Scope: High-consequence identity matching
KPIs: False-match and missed-match rates
Model: Assurance-led implementation
Dependency: Clinical and legal review

Migration and system consolidation

Resolve duplicates before records are migrated into a new CRM, ERP, MDM, lakehouse or customer-data platform.

Scope: Pre-migration matching and reconciliation
KPIs: Exceptions, reconciliation closure
Model: Project delivery
Dependency: Source extracts and cutover plan

Fraud and abuse investigation

Link accounts, devices, addresses and behaviours to surface possible networks while retaining explainability and investigator oversight.

Scope: Relationship and network resolution
KPIs: Alert quality, analyst review burden
Model: Specialist advisory
Dependency: Risk and fairness controls

Citizen or constituent services

Improve service coordination across programmes without treating every similar record as the same person or exceeding permitted data use.

Scope: Cross-programme identity linkage
KPIs: Match confidence and correction rate
Model: Phased programme
Dependency: Statutory and policy authority
Capabilities

Identity resolution capabilities from evidence to operations

Entity definition, source discovery and data profiling

Define the real-world entity, relationships and permitted uses; inventory source systems; profile identifiers and attributes; measure completeness, uniqueness and conflict; map lineage; and identify source-authority assumptions. Deliverables can include an entity scope, source map, profile findings, issue register and assessment baseline.

Matching strategy and model design

Design deterministic rules, standardisation, tokenisation, phonetic and fuzzy comparisons, probabilistic scores, blocking and candidate generation, thresholds, confidence bands and clerical review. Technology choices may include MDM tools, cloud services, graph capabilities, data platforms or custom libraries. Rules are tested against labelled examples and business risk.

Survivorship, golden records and relationship management

Define attribute-level survivorship, source precedence, recency, validity, trust scores, merge and unmerge controls, crosswalks, householding, organisational hierarchies and relationship types. Outputs include a golden-record model, rulebook, exception paths and ownership matrix.

Implementation, tuning and integration

Support configuration, pipelines, APIs, batch and event-driven processing, reference data, test harnesses, reconciliation, migration, performance tuning and downstream publication. Implementation is adapted to existing platforms and release controls rather than assuming a single vendor product.

Governance, stewardship and monitoring

Establish accountable data owners, steward queues, reason codes, rule approval, model change control, quality thresholds, false-match investigation, access control, audit evidence and operational reporting. Managed support can be scoped for ongoing tuning and issue management.

Deliverables

Service deliverables aligned to the agreed scope

Deliverables are selected according to the entity type, risk, maturity, technology environment and delivery stage.

Representative identity resolution deliverables
DeliverableWhat it includesFormatStageClient input requiredPrimary owner
Identity resolution assessmentSource inventory, profiling, duplicate patterns, risks, readiness and recommendationsReport and findings registerDiscoveryData samples, schemas, issue logsData lead
Entity and identity modelEntity boundaries, identifiers, relationships, golden-record and crosswalk designModel and design specificationDesignBusiness definitions and use casesData architect
Matching and survivorship rulebookRules, weights, thresholds, precedence, reason codes and exception treatmentControlled rule catalogueDesign and tuningRisk tolerance and labelled examplesData owner
Solution architectureComponents, flows, integration, security, environments, monitoring and non-functional needsArchitecture packDesignPlatform constraints and standardsArchitecture owner
Validation and test packTest datasets, precision and recall analysis, edge cases, reconciliation and acceptance evidenceTest cases and resultsImplementationApproved acceptance criteriaQA lead
Stewardship and operating modelRoles, queues, SLAs, escalation, rule change, audit and reportingProcedures and RACITransitionNamed owners and capacityService owner
KPI and monitoring frameworkQuality, model, operational and control metrics with thresholds and review cadenceDashboard specificationOperateBaseline and reporting platformService manager

Define the deliverables your identity problem requires

Scope an assessment, design, implementation, assurance review or managed operating service.

Request a Consultation
Service process

How Dataconsultant delivers identity resolution

The sequence is adapted to scope, but each stage produces an explicit decision or controlled output.

Discovery and alignment

Objective: Confirm entity types, business outcomes, risks and stakeholders.

Output: Scope, decision log and evidence request.

Source and quality assessment

Objective: Understand identifiers, duplication, conflicts and constraints.

Output: Profile findings and source map.

Match strategy design

Objective: Define candidate generation, comparison, scoring and thresholds.

Output: Matching rulebook and test plan.

Golden-record and control design

Objective: Establish survivorship, stewardship, privacy and audit controls.

Output: Target model, workflow and RACI.

Build, configure and tune

Objective: Implement matching, integration and exception handling.

Output: Configured capability and tuning evidence.

Validate and transition

Objective: Test accuracy, reconcile outputs and prepare operations.

Output: Acceptance pack, training and monitoring baseline.

Technology and frameworks

Platforms, methods, standards and control references

Technology is selected around the identity use case, data volumes, latency, risk, explainability, integration and operating capability.

Technology patterns

  • MDM platforms
  • Customer data platforms
  • Cloud data platforms
  • Graph databases
  • Data quality tools
  • Streaming and APIs
  • Search indexes
  • Python and SQL

Matching methods

  • Deterministic linkage
  • Probabilistic linkage
  • Fuzzy comparison
  • Phonetic matching
  • Tokenisation
  • Blocking
  • Clustering
  • Human review

Relevant references

  • DAMA-DMBOK
  • ISO 8000 principles
  • ISO 27001 controls
  • Privacy-by-design principles
  • NIST privacy and security guidance
  • Sector-specific requirements
  • Internal data standards
  • Records management policies

Assess your existing platform before adding another tool

Dataconsultant can evaluate whether current MDM, CRM, cloud or data-platform capabilities can support the required identity controls.

Request a Consultation
Engagement models

Flexible ways to engage

Identity resolution engagement models
ModelBest suited toTypical scopeClient responsibilityCommercial basis
Focused assessmentUnclear scale, risk or technology directionProfiling, findings, options and roadmapProvide samples, stakeholders and contextFixed scope or time-boxed
Design engagementTeams preparing procurement or implementationRules, architecture, operating model and test designApprove definitions and thresholdsMilestone-based
Implementation supportExisting platform or engineering programmeConfiguration, pipelines, tuning, testing and transitionPlatform access, delivery governance and acceptanceProject or capacity-based
Independent assuranceHigh-risk or vendor-led programmesDesign review, model testing, controls and evidenceProvide artefacts and remediation ownershipDefined review scope
Managed optimisationOperational matching requiring sustained tuningMonitoring, exception analysis, rule changes and reportingRetain accountable owners and approve changesRecurring service
Illustrative examples

How identity decisions may be handled

These examples explain decision patterns only and do not represent actual client results.

Exact identifier, conflicting address

Two records share a validated account identifier but contain different addresses. The records may be linked while address survivorship follows recency, validation status and source authority. The previous value remains traceable.

Similar name, shared household

Two people have similar names and the same address but different dates of birth and contact details. The design should preserve separate person identities while optionally creating a governed household relationship.

High score near a review threshold

A probabilistic score indicates a likely match, but the consequence of error is high. The candidate is routed to a steward with reason codes and supporting evidence instead of being merged automatically.

Expected outcomes and KPIs

Measure both match quality and operational control

Targets should be based on a documented baseline and the cost of false matches, missed matches and manual review.

Representative identity resolution measures
KPIWhat it indicatesImportant interpretation
PrecisionProportion of predicted matches that are correctHigh precision matters where false linking has material consequences.
RecallProportion of true matches successfully identifiedImproving recall can increase false positives without careful tuning.
Duplicate rateResidual duplicate entities or recordsRequires a consistent definition and comparable source population.
Manual review rateVolume routed to human decisionShould balance risk, cost and steward capacity.
Exception ageingTime unresolved candidates remain in queuesDepends on ownership, staffing and escalation design.
Unmerge or correction rateIdentity decisions later reversed or correctedA useful control signal, not proof of total model accuracy.
Rule-change stabilityImpact of model or threshold changesRequires versioning, regression tests and monitored deployment.
Downstream adoptionUse of trusted identity outputs by approved consumersMust be measured alongside privacy, consent and access compliance.
Pricing and cost factors

What influences identity resolution cost

A reliable estimate requires initial scoping because data complexity and control requirements vary significantly.

Data scope

Number of sources, entities, records, languages, jurisdictions, attributes and historical data volumes.

Match complexity

Identifier quality, relationship modelling, probabilistic methods, labelled data and threshold tuning.

Technology environment

Existing licences, platform constraints, integrations, real-time needs, performance and deployment environments.

Governance and assurance

Privacy review, security controls, audit evidence, stewardship workflows, testing depth and regulatory obligations.

Request a scoped estimate

Share the entity type, source landscape, target platform, key risks and expected delivery stage.

Request a Consultation
Why Dataconsultant

A business, data and control-led approach

Business definitions before algorithms

We clarify what an entity means, why records must be linked and what harm an incorrect decision could cause before selecting methods.

Evidence-conscious delivery

Recommendations distinguish observed findings, assumptions, illustrative examples and matters requiring specialist legal, privacy or security review.

Vendor-neutral implementation support

We can work with internal teams, platform vendors and systems integrators while maintaining clear ownership, acceptance and assurance boundaries.

Discuss your identity resolution requirement

Start with a practical review of the business problem, data evidence, delivery options and next decisions.

Request a Consultation
Security, quality, privacy and compliance

Controls must reflect the consequence of identity decisions

Security

Data classification, least-privilege access, secure transfer, environment separation, logging, secrets management and controlled administrative actions.

Quality

Profiling, standardisation, validation, representative test data, precision and recall checks, regression testing and reconciliation.

Privacy

Purpose limitation, lawful basis, consent, data minimisation, transparency, retention, sensitive attributes and correction processes.

Compliance

Applicable sector, contractual, records-management, residency and audit requirements, validated with authorised specialists where necessary.

Delivery environment

Technology ecosystems identity resolution may connect

Identity resolution typically operates between operational systems, enterprise data platforms and governed consuming applications.

Source and engagement systems

CRM, ERP, ecommerce, support, billing, identity and access, case management, clinical, policy, claims and line-of-business applications.

Data and identity platforms

MDM, CDP, data quality, integration, lakehouse, warehouse, graph, streaming, search and metadata platforms.

Downstream consumers

Customer service, analytics, finance, risk, fraud, marketing, operations, regulatory reporting and authorised AI or decision-support use cases.

Customer perspectives

Representative identity resolution testimonials

These service-specific testimonials illustrate the types of delivery experience customers may value. They are not presented as independently verified performance claims.

★★★★★
“The team helped us move from vague duplicate concerns to a documented identity model, clear match reasons and an exception process our operations staff could actually use. Communication stayed practical, and the design decisions were explained without forcing us into a particular platform.”
Customer Data DirectorRetail and Ecommerce
★★★★★
“Our supplier records were spread across procurement and finance systems with inconsistent legal names and identifiers. Dataconsultant structured the assessment carefully, highlighted where external validation was needed, and produced survivorship and hierarchy rules that our internal team could review and maintain.”
Head of Procurement OperationsManufacturing
★★★★★
“The strongest part of the engagement was the attention to false-match risk. The consultants worked with privacy, service and data teams, created review thresholds, and made sure every automated decision could be traced back to understandable evidence and an approved rule.”
Data Governance LeadHealthcare Services
★★★★★
“We needed independent support during a CRM consolidation. The team reviewed source quality, challenged optimistic vendor assumptions, and gave us a realistic test and reconciliation plan. Revision requests were handled professionally, and ownership remained clear throughout the delivery.”
Technology Programme ManagerProfessional Services
★★★★★
“Dataconsultant translated complex matching concepts into decisions our risk and operations leaders could make. The resulting rulebook, stewardship workflow and monitoring measures gave us a more controlled basis for improving account linkage without treating similarity as certainty.”
Risk Analytics ManagerFinancial Services
★★★★★
“The engagement covered more than an algorithm. It connected data profiling, architecture, security, privacy and service operations. Knowledge transfer was thorough, documentation was organised, and our engineers and data stewards finished with a shared understanding of how the capability should be governed.”
Chief Data OfficerPublic Sector
Frequently asked questions

Identity resolution FAQs

What is the difference between identity resolution and data deduplication?

Deduplication usually focuses on finding and removing duplicate records within a defined dataset. Identity resolution is broader: it determines whether records across systems refer to the same entity, maintains links and crosswalks, applies survivorship, manages ambiguous cases, records provenance and supports ongoing governance.

What entity types can be resolved?

Common entity types include customers, patients, citizens, members, employees, suppliers, counterparties, organisations, products, locations, assets, accounts and devices. Each requires its own definition, identifiers, relationships, risk thresholds and governance.

Do we need an MDM platform?

Not always. Some needs can be addressed within an existing CRM, cloud data platform, CDP, warehouse or integration architecture. An MDM platform may be appropriate when persistent golden records, cross-domain governance, workflows, hierarchies and broad integration are required. The decision should follow assessment.

How do deterministic and probabilistic matching differ?

Deterministic matching uses explicit rules, such as equality on a trusted identifier. Probabilistic matching combines evidence from multiple attributes to estimate whether records belong to the same entity. Many solutions use both, with thresholds and manual review based on risk.

How are false matches and missed matches controlled?

Controls can include representative labelled data, precision and recall measures, confidence bands, stricter thresholds for high-consequence use cases, reason codes, manual review, regression tests, monitoring, correction workflows and controlled rule changes.

What data is required to start?

Useful inputs include source schemas, representative data samples, identifier definitions, data-quality reports, duplicate examples, business processes, privacy and security requirements, existing rules, platform architecture, issue logs and access to accountable business and technical owners.

How long does an identity resolution engagement take?

There is no reliable fixed duration without discovery. Timing depends on the number and quality of sources, entity complexity, data volume, labelled examples, platform readiness, integration scope, review cycles, privacy requirements, testing depth and whether implementation is included.

How is identity resolution pricing calculated?

Pricing is influenced by source count, data volume, entity types, matching complexity, technology, integrations, real-time requirements, governance design, testing, security and privacy review, stewardship workflows, deployment support and the selected engagement model.

Can Dataconsultant work with our existing vendor?

Yes. Dataconsultant can support requirements, design, implementation, tuning, testing, assurance and operating-model work alongside internal teams, platform vendors and systems integrators. Responsibilities, access, acceptance criteria and escalation routes should be documented.

Can identity resolution support fraud detection?

It can support fraud and abuse analysis by linking accounts, devices, addresses, payment instruments and other relationships. Such use requires careful governance, explainability, access control, fairness consideration, investigator oversight and validation appropriate to the decision impact.

How are privacy and consent handled?

The design should document purpose, lawful basis, consent and preference treatment, sensitive attributes, access, retention, residency, correction and downstream use. Legal and regulatory interpretations must be validated by authorised advisers for the relevant jurisdictions and sector.

What happens when the system is uncertain?

Uncertain candidates can remain unlinked, be provisionally linked, or be routed to a steward depending on policy and risk. The workflow should expose evidence, confidence, reason codes and consequences so that reviewers can make consistent decisions.

How is a golden record maintained over time?

Ongoing maintenance requires source crosswalks, incremental matching, survivorship, versioned rules, monitoring, stewardship, merge and unmerge controls, data-quality remediation, ownership and controlled distribution to downstream systems.

Can the service include managed support?

Yes. Managed support can cover monitoring, exception analysis, rule tuning, quality reporting, steward assistance, release assurance and continuous improvement. Accountable business ownership and approval authority should remain clearly assigned.

What are the main limitations of identity resolution?

No method can guarantee perfect identity decisions when evidence is incomplete, inaccurate, outdated or legally unavailable. Accuracy also depends on definitions, representative testing, threshold choices and operational controls. Identity resolution does not replace source remediation, legal advice, formal audit or specialist security testing.