Strong data modeling skills, specifically with Databricks.
5+ years in data modeling or architecture, particularly with unstructured/semi-structured data.
Experience in Knowledge Management or related fields, with insight into document-based data.
Hands-on experience with data catalogs, particularly Databricks Unity Catalog.
Ability to define data domains and delineate product boundaries in complex organizations.
Knowledge of metadata management practices and frameworks for classification.
Understanding of data security and its integration within a Lakehouse environment.
Responsibilities
Design logical and physical data models for unstructured and semi-structured content.
Define ownership and boundaries for reusable data products.
Establish consistent metadata and tagging standards across knowledge sources.
Assign classifications to data products in compliance with data governance and privacy requirements.
Document and maintain data products in Databricks Unity Catalog, including metadata.
Collaborate with data engineers to align pipeline practices with modeled domains.
Align data product design with the needs of various stakeholders, including Knowledge and Research teams.
Benefits
Remote work flexibility in Canada.
Opportunity to shape the foundational aspects of a new KM Data Platform team.
Engagement in innovative projects involving unstructured data handling.
Collaboration with multi-disciplinary teams, enhancing professional development.
Full Job Description
Role: Data Modeler Location: Remote Canada Employment type: Fulltime
Job Overview:
The Senior Data Modeler will design and govern the data architecture for unstructured knowledge assets across Knowledge Management (KM) ecosystem.
This role bridges data engineering discipline with KM domain expertise, translating raw unstructured content (documents, case files, informal knowledge captures, chat/email extracts, etc.) into well-defined, discoverable, and secure data products within Databricks Unity Catalog.
This is a foundational hire for a newly formed KM Data Platform team supporting broader Knowledge and Research Systems strategy.
Key Responsibilities:
Design logical and physical data models for unstructured and semi-structured content (documents, case artifacts, K-Slices, extracted knowledge fragments, metadata records) originating from KM pipelines such as case mining and informal knowledge capture workflows.
Define domain boundaries and ownership for data products - determining what constitutes a discrete, reusable data product versus a raw or intermediate asset.
Establish metadata standards and tagging taxonomies (content type, practice/domain, provenance, confidentiality, freshness, lineage) to ensure consistent classification across knowledge sources.
Assign and enforce security and sensitivity classifications on data products in line with firm data governance, privacy, and legal/risk requirements.
Register, document, and maintain data products in Databricks Unity Catalog, including schemas, access grants, lineage, and catalog-level metadata.
Partner with data engineers building Databricks pipelines to ensure ingestion, transformation, and storage patterns align to the modeled domain structure.
Collaborate with Knowledge Products, Research Products, and Architecture/Data/Technology stakeholders to align data product design with downstream consumption needs (e.g., surfacing in Sage/Glean, AI agent retrieval).
Support privacy and legal review processes by ensuring data products are classified and documented to enable timely sign-off.
Establish and document repeatable modeling standards/playbooks so future data products can be onboarded consistently as the KM platform scales.
Qualifications
Must have skills:
Data Modeling (Strong), Databricks.
5+ years of experience in data modeling, data architecture, or information architecture, with meaningful exposure to unstructured or semi-structured data (not purely relational/transactional modeling).
Direct experience working in or adjacent to Knowledge Management, content management, or enterprise search domain - understands how documents, case files, or knowledge artifacts differ from standard transactional data.
Hands-on experience with a modern data catalog; Databricks Unity Catalog experience strongly preferred.
Demonstrated ability to define data domains and data product boundaries in a large, multi-stakeholder organization.
Practical knowledge of metadata management: tagging schemas, taxonomies, controlled vocabularies, or ontology design.
Understanding of data security/sensitivity classification frameworks and how they map to access control in a Lakehouse environment.
Experience partnering with data engineering teams on ingestion and pipeline design (not required to write production pipeline code, but must speak the language).
Strong written and verbal communication skills; able to translate technical modeling decisions into business-readable rationale for KM stakeholders and governance reviewers.
Good to have skills:
BI Schema Design - General Experience.
Experience with enterprise knowledge platforms (e.g., Glean, SharePoint, ServiceNow) or AI-powered retrieval systems.
Familiarity with Databricks Delta Lake, Delta Sharing, or Lakehouse Federation.
Prior experience in professional services, consulting, or a similar document/case-intensive knowledge environment.
Exposure to Legal/Risk/Privacy review processes for data classification and access approvals.
Background in library science, information science, or applied ontology is a plus but not required. Success Metrics (First 6-12 Months)
Domain model and metadata taxonomy defined and adopted for at least one major KM data product line (e.g., case mining outputs, informal knowledge K-Slices).
Data products registered and discoverable in Unity Catalog with correct security classifications applied.
Documented, repeatable modeling standard that engineering and future modelers can apply without re-litigating domain boundaries each time.
Reduced turnaround time on privacy/legal classification reviews due to upfront, consistent metadata and tagging.