Noah Antle Avatar

Golden Lake is Entrada’s Databricks-native MDM and entity resolution accelerator, built to resolve your most valuable entities without ever moving data out of Unity Catalog. Master Data Management (MDM) is a constant source of tension for data-driven organizations. There is a crucial need to obtain a unified, 360-degree perspective of core entities within their data, like customers, vendors, patients, products, etc. but achieving this clarity has historically required potentially-risky movement of sensitive data to and from a data lake and through a legacy MDM tool which requires additional infrastructure, expertise, and costs. 

The traditional MDM approach can lead to a variety of problems for an organization. Extracting and moving data across systems degrades governance, creates latency, and relies on infrastructure which wasn’t designed for today’s massive data volumes. The architecture of MDM needs to evolve.

At Entrada, we believe that modernizing MDM means bringing the underlying entity resolution engine and user experience directly to your data stored in Unity Catalog. This is why we built Golden Lake: a Databricks-native MDM and Entity Resolution solution accelerator. Golden Lake is designed to unify duplicate, fragmented, and messy records into a single “golden record” per real-world entity and runs entirely on your existing Databricks compute with all of the governance provided by Unity Catalog; data doesn’t need to leave your Databricks environment at any point and no third-party MDM solution is required, resulting in a fast, governable, auditable, and streamlined MDM solution.

The Transition to Databricks-Native MDM

Legacy MDM platforms tend to be black boxes which introduce fragmented processes, external licensing, and manual validation steps that are not always scalable and don’t always cooperate with a Databricks environment. Golden Lake replaces legacy bottlenecks with a platform built for speed, leaning on the governance features baked into Databricks and the Unity Catalog. By building Golden Lake on Spark for compute, Delta Lake for storage, and Unity Catalog for security and governance, the data running through the entity resolution pipeline never needs to leave your Databricks environment. It is able to seamlessly transition from raw bronze data, to standardized silver tables, and finally into gold master records, relationships, and metadata while maintaining lineage and access controls.

The Golden Lake Architecture

Golden Lake is built on two core components: a scalable PySpark pipeline engine for entity resolution, source linkage, and metadata creation and logging, and a no-code/low-code Streamlit interface hosted via Databricks Apps.

The MDM Engine

The backend engine is deployed as a Python wheel to a Unity Catalog Volume and executes as a multi-task Databricks Job on an ephemeral job cluster (as of the time of development, the underlying fuzzy-matching library used for entity resolution, Splink, is not compatible with Serverless compute). Splink, a probabilistic record linkage Python library) handles complex fuzzy matching at massive scale.

The pipeline runs sequentially through optimized stages:

Pipeline StageFunction
Change DetectionChecks source tables for new/changed rows; skips processing if unchanged to save compute.
StandardizationMaps heterogeneous columns to a unified schema and applies text/whitespace normalization.
ValidationApplies user-defined business rules (null/format checks) and flags invalid records.
MatchingUses configurable blocking rules and fuzzy comparison algorithms to cluster records above a match threshold (configurable by the user).
Golden RecordMerges clustered records into a single best-version entity using layered survivorship strategies.
RelationshipsResolves cross-entity references and writes a comprehensive relationship graph.

The Golden Lake UI (Databricks App)

Golden Lake, built using Databricks Apps, was designed to be a frictionless user experience for data stewards and business users, even if they don’t possess a highly-technical background or extensive knowledge of the underlying data. 

image

Deployed via Declarative Automation Bundles (DABs), the app provides a centralized hub for end users to define entities, add source tables to existing entities, configure the entity resolution pipeline, and trigger/schedule/monitor resolution jobs. The pipeline leverages a config-driven architecture where business users configure rules via the app’s UI, instantly updating a YAML file stored in a Unity Catalog Volume. The engine interprets this config file when it runs, meaning a user does not need to write or modify any code to configure and run an entity resolution job.

AI-Assisted Configuration with Human-in-the-Loop Governance

One of the biggest hurdles in a traditional MDM implementation is the effort and technical knowledge required to profile the data which will be traveling through the entity resolution pipeline and configure the matching rules that tell Splink how to perform the matching. Golden Lake eases this difficulty by integrating Foundation Model endpoints directly into the business rule configuration workflow.

  • Intelligent Mapping and Rules: The AI profiles your source and target schemas to recommend optimal column mappings, data quality validation rules, and the most effective blocking rules based on your actual data.
image 1
  • Optimal Comparisons: Golden Lake can recommend which string similarity algorithms (e.g., Jaro-Winkler, Levenshtein) and confidence thresholds for these algorithms to apply on a field-by-field basis.
image 2
  • Alias Detection: The engine identifies known synonyms (e.g., “IBM” vs. “International Business Machines”) and publishes them to a real-time lookup table that can be used in downstream queries.
image 3

For MDM experts or those who want to experiment, all of these settings are manually configurable, but AI suggestions are always available to streamline the setup process and get a successful entity resolution job up and running.

However, automation should never be prioritized over accuracy. AI accelerates the configuration steps of the MDM engine, but Golden Lake enforces strict governance through a Human-in-the-Loop Match Review system, where data stewards can review uncertain matches (based on your configured confidence thresholds) that are routed to a review queue in Unity Catalog. Stewards can use the UI to approve, reject, or manually override cluster memberships, and these manual overrides always take the highest precedence in Golden Lake’s survivorship model, ensuring that human expertise always acts as a final guardrail.

image 4

Genie-Assisted Exploration of Outputs

Golden Lake includes a Genie Data Assistant tab containing an embedded Genie Agent, which provides a natural language interface over your source, silver, and gold MDM tables without needing to leave the app. This helps to bridge the gap between the underlying entity resolution engine and data analysis, meaning users can ask ad-hoc analytical questions (e.g. “how many duplicate vendors were resolved last run?” or “show me all sources mapped to the customer entity”) and receive instant SQL-backed responses from the Genie Agent. Additionally, the Genie Agent automatically syncs with any new tables that get created and saved into the MDM Unity Catalog schema, ensuring that new entities that get created and resolved by the pipeline will be immediately available to use for Genie-based insights without any manual configuration required.

image 5

Databricks-Native MDM for the Future of Analytics

MDM solutions deliver value across an entire organization, across all technical and operational domains. Golden Lake drastically reduces the wait times, governance risks, and infrastructure headaches associated with legacy tools by keeping all source and output data within your Databricks environment and running jobs using usage-based, ephemeral job clusters for efficient multi-task job runs.

Golden Lake provides a highly governable foundation, with master data accurately resolved, linked, and enriched within Unity Catalog, leaving your organization in a strong position to support downstream workflows like traditional BI or AI-driven search and discovery.

The issue of master data management needn’t be an architectural roadblock. Leveraging the combined capabilities of Databricks Apps and scalable Spark pipelines, Golden Lake develops a stable, native operating model to bring together, govern, and then use critical sensitive entity data.

GET IN TOUCH

Millions of users worldwide trust Entrada

For all inquiries including new business or to hear more about our services, please get in touch. We’d love to help you maximize your Databricks experience.