Senior Data Scientist Role SummaryThe Senior Data Scientist will be responsible for analyzing, cleansing, reconciling, and statistically validating complex datasets related to wildlife trade, permits, applicants, species, transactions, shipments, and historical records. The role will focus heavily on data quality, entity resolution, anomaly detection, taxonomy reconciliation, and the development of reliable, standardized datasets for downstream analytics and reporting. The ideal candidate will have strong experience working with large, complex, and legacy datasets and will be comfortable developing reproducible analytical solutions using Python, SQL, and preferably R.
Key Responsibilities- Analyze and cleanse complex structured and unstructured datasets from multiple sources.
- Identify malformed, irrelevant, incomplete, and corrupted free-text records and extract usable structured information.
- Standardize and map data values to established controlled vocabularies and master-data standards.
- Reconcile species identifiers, scientific names, common names, and related taxonomy information.
- Resolve synonyms, misspellings, deprecated names, taxonomy inconsistencies, and other data discrepancies.
- Standardize units of measure, quantities, country codes, source codes, and other data attributes.
- Perform deterministic and probabilistic record linkage across multiple datasets.
- Apply fuzzy matching and similarity techniques to identify related records.
- Identify duplicate applicants, permits, transactions, shipments, and other business entities.
- Develop reversible and auditable entity-resolution and entity-merging methodologies.
- Design and implement statistical methods for anomaly and outlier detection.
- Analyze cross-source consistency, completeness, accuracy, and reliability.
- Develop data-quality rules, validation checks, metrics, and scorecards.
- Establish measurable data-quality thresholds and monitor data-quality trends.
- Measure model and data-matching performance using labeled datasets, including precision and recall.
- Support the validation and certification of Gold-layer datasets.
- Develop reproducible Python and/or R analytical code and SQL queries.
- Document data-cleansing, reconciliation, matching, validation, and quality-assurance methodologies.
- Collaborate with data engineers, analysts, subject-matter experts, and project stakeholders to resolve complex data-quality issues.
- Provide recommendations for improving data standards, governance, and long-term data integrity.
Required Qualifications- 7-10+ years of experience in data science, data quality, data analytics, or a related field.
- Strong experience working with complex, large-scale, and/or legacy datasets.
- Strong programming experience with Python and SQL.
- Experience with R is preferred.
- Hands-on experience with entity resolution, record linkage, duplicate detection, or master-data management.
- Experience implementing deterministic, fuzzy, and probabilistic matching techniques.
- Strong understanding of statistical analysis, anomaly detection, and outlier identification.
- Experience developing data-quality rules, validation frameworks, metrics, and scorecards.
- Experience working with controlled vocabularies, reference data, and standardized coding systems.
- Strong experience with data cleansing, transformation, normalization, and reconciliation.
- Experience evaluating analytical results using labeled datasets and measures such as precision and recall.
- Ability to develop reproducible, maintainable, and well-documented analytical code.
- Strong analytical, problem-solving, and communication skills.
Preferred Qualifications- Experience working with scientific, environmental, biological, ecological, or taxonomy-related datasets.
- Knowledge of wildlife trade data, species identification, scientific nomenclature, or biological taxonomies.
- Experience resolving scientific names, synonyms, deprecated classifications, and taxonomy mismatches.
- Experience with large-scale legacy-data remediation and historical data conversion.
- Experience supporting data certification or Gold-layer data products.
- Familiarity with data governance, master-data management, and enterprise data-quality frameworks.
- Experience working with multiple heterogeneous data sources and developing cross-source validation methodologies.
- Experience in government, regulatory, environmental, or scientific data environments is a plus.
Does this opportunity sound like a fit for you? If so, join our talent community and click to apply now!!