
A master-data model is a decision system expressed as structure. It decides what counts as an entity, which identifiers persist, which relationships are allowed and which contradictions must stop a transaction or merely warn a steward.
Audience, prerequisites and outcomes
For data modelers, domain architects and EBX stewards. Complete Parts 1–2 and obtain model-author permission in a non-production EBX 6.x environment. You will create normalized customer and product tables, controlled vocabularies, provenance fields and validation rules, then test them with synthetic data.
Version note: the Data Model Assistant and advanced controls evolve. The concepts—tables, fields, primary/foreign keys, facets and constraints—are stable; verify UI labels against your installed 6.x release.
Model the entity, not the source screen
Use durable master keys unrelated to CRM or ERP identifiers. Preserve every external key in an identifier/source table.
Core tables
| table | primary key | important fields |
|---|---|---|
| Party | party_id | partytype, displayname, status |
| Customer | customer_id | partyid, lifecyclestatus, market |
| Product | product_id | displayname, lifecyclestatus, categorycode, brandcode |
| ExternalIdentifier | identifier_id | entitytype, masterid, sourcesystem, sourceid, valid_from/to |
| Address | address_id | partyid, type, lines, locality, region, postalcode, country_code |
| ProductIdentifier | identifier_id | product_id, scheme, value, market |
| Category | category_code | name, parentcategorycode |
| SourceRecord | sourcerecordid | sourcesystem, sourcekey, receivedat, payloadhash |
Separate Party from Customer if individuals and organizations share identity behavior. Do not force this pattern when the domain does not need it.
Constraints as governed policy
EBX supports XML Schema facets plus extended and programmatic constraints. Choose severity deliberately:
- blocking error: impossible or unsafe state, such as an unknown country code;
- error routed for correction: invalid operational value;
- warning: plausible but suspicious, such as missing optional description;
- information: stewardship prompt.
Suggested controls:
| field/rule | control |
|---|---|
partyid, productid | mandatory; generated or governed pattern |
| source + source_id | unique together |
| country_code | foreign key to controlled reference |
| normalized representation plus format check; never identity alone | |
| product status | enumeration: Draft, Active, Suspended, Retired |
| category | foreign key; prevent invalid self-parenting |
| effective dates | validto >= validfrom when present |
| AI description | exclude secrets/PII; steward approval before consumption |
Hands-on in EBX
Establish a recoverable model workspace
Action: Confirm model-author permission, record the installed EBX 6.x version, create a child model workspace/dataspace, and capture a snapshot or export. Expected result: northstar_master changes are isolated from production data. Validate: compare the baseline model version and lab target identifiers. Stop/recover: if changes appear in a production-connected dataset, stop, revert to the captured baseline, and recreate the exercise in isolation.
Create groups and reference tables
Action: Add form/model groups Identity, Contact, Classification, Lifecycle, and Provenance. Create Country, CustomerStatus, ProductStatus, Category, and SourceSystem first, each with a stable key, label, lifecycle status, and any required effective dates. Load the workshop codes, including US, before dependent tables. Expected result: reference records validate independently. Validate: duplicate keys, blank labels, unknown parent category, and self-parent category produce documented outcomes. Stop/recover: do not create dependent foreign keys while reference keys or hierarchy rules are unstable.
Create master and identifier tables
Action: Create Party, Customer, Product, ExternalIdentifier, Address, and ProductIdentifier with the documented keys and cardinalities. Keep partyid and productid independent of source IDs; make source-system plus source ID unique in the identifier structure. Expected result: the supplied customer and product samples load with durable master identifiers. Validate: change an external ID and prove the master ID remains stable; attempt duplicate (CRM,C-1007). Stop/recover: if a source key became the master primary key, revise the model before loading more data rather than attempting an in-place production key change.
Add relationships and quarantine policy
Action: Add foreign keys from customer/address to party, product to category/status, and identifier rows to their master/source records. For each missing reference, document whether the landing import stops, the record quarantines, or mastering blocks; do not use one behavior implicitly everywhere. Expected result: valid relationships resolve and country_code=USA or category UNKNOWN follows the chosen controlled path. Validate: inspect constraint messages and the quarantine record's rule, source key, batch, severity, and owner. Stop/recover: if invalid references enter approved master data, block promotion and remove only the affected lab records after preserving evidence.
Configure declarative constraints first
Action: Add mandatory cardinality, length, patterns, enumerations, uniqueness, and effective-date rules with risk-appropriate severity. Use an advanced or programmatic constraint only for policy that cannot be expressed declaratively, and document its code/version and execution scope. Expected result: deliberate failures produce specific actionable messages. Validate: test blank durable keys, duplicate source ID, invalid email syntax, unknown country/category, and validto before validfrom. Stop/recover: if a programmatic rule errors, times out, or blocks unrelated records, disable that lab rule and replace it with a tested implementation before publication.
Protect provenance and system audit fields
Action: Add createdat, createdby, updatedat, updatedby, sourcerecordid, and governance_status; configure services/permissions so users cannot forge system-managed values. Expected result: create/update operations populate accountable history while source linkage remains intact. Validate: attempt manual overwrite with a steward test account and verify denial or controlled service behavior. Stop/recover: treat writable audit fields as a release blocker and correct permissions before model promotion.
Publish and validate in the lab
Action: validate the semantic model, publish a new lab model version using the installed-release workflow, create or refresh northstar_lab, and load both valid and deliberate-failure rows. Expected result: valid rows persist, deliberate failures match the planned severities, and the validation report identifies each failing rule. Validate: reconcile expected versus observed results row by row and save the model version/report. Stop/recover: keep the previous model active if publication or refresh fails; restore the lab snapshot/export and resolve schema migration issues before retrying.
Synthetic records to enter
party_id,party_type,display_name,status
PTY-00017,PERSON,Amina Rahman,ACTIVE
PTY-00018,ORGANIZATION,Blue Peak Retail,ACTIVEproduct_id,display_name,lifecycle_status,category_code,brand_code
PRD-00200,Alpine Pack 40L,ACTIVE,PACKS,NORTHSTAR
PRD-00201,Storm Shell,ACTIVE,OUTERWEAR,NORTHSTARAdd deliberate failures: country_code=USA when the code list expects US; duplicate (CRM,C-1007); product category UNKNOWN; and a validity end before its start. Record the resulting severity and message.
AI-aware fields
Do not create a generic ai_text dumping ground. Model intended derivatives explicitly:
approved_summary: human-approved factual description;search_keywords: governed discovery terms;embedding_policy: whether and for which purposes embedding is allowed;content_language: language code;sensitivity: classification;effectiveatandrecordversion: reproducibility anchors.
Generated descriptions remain proposed content until reviewed. Store model/prompt provenance outside the canonical business attribute or in a linked derivation table.
Exercise: change without breaking consumers
Exercise: evolve localization without breaking identity
Action: copy the two sample products into the lab and capture their product_id, current names, model version, and consumer response. Compare repeated locale fields, a ProductName child table keyed by product/language/market, and a separate localization dataset against cardinality, governance, fallback, query, and release requirements. Record the selected design and rejected alternatives. Expected result: one explicit design with a deterministic fallback such as market-language → language → approved default.
Action: implement the chosen design in the child model version, migrate names for both products, validate uniqueness and required-language rules, then exercise the same consumer query. Expected result: localized names resolve according to fallback and both original product_id values remain unchanged. Validate: compare before/after identifiers and test missing locale, duplicate locale/market, retired product, and rollback. Stop/recover: if identifiers change or the existing contract breaks, do not promote; restore the baseline and introduce a backward-compatible view or staged consumer migration.
Validation checklist
- Not completed: Master keys do not depend on source-system keys.
- Not completed: External identifiers are unique within source and retain validity.
- Not completed: Reference values use foreign keys or controlled enumeration.
- Not completed: Constraint messages tell stewards how to act.
- Not completed: Severity matches business risk.
- Not completed: AI-derived text is distinguishable from approved facts.
- Not completed: Model change preserves identifiers and consumer compatibility.
- Not completed: The validation report contains expected deliberate failures.
Failure modes
Customer equals email. Emails change, can be shared and can be mistyped.
Every source attribute becomes canonical. A master model is curated semantics, not a union of source schemas.
All constraints block. Landing cannot preserve imperfect evidence and stewards cannot see the real quality problem.
Free-text reference values. “US,” “USA” and “United States” become three facts.
Overusing programmatic controls. Hidden code is harder for stewards to understand than declarative structure.
Takeaway and next
A strong model separates identity, attributes, references and evidence. It makes unsafe states difficult while leaving uncertainty visible. Part 4 brings source records into the model and turns validation results into quality measurements.

Historical comments from Datanizant
No public comments on this article
No approved public comments were included in the WordPress export for this article.