Job Summary
We are seeking a detail-oriented and technical Data Analyst to lead synthetic data curation,
relational modeling, and dataset validation for an AI-driven, multi-industry demo platform. In this
role, you will design synthetic data models across four distinct domains (Financial, Healthcare,
Manufacturing, and Retail). You will write programmatic generation scripts to produce realistic,
non-PII datasets, load them into an Oracle Autonomous Database (ADB), and validate dataset
coverage to ensure accurate Natural-Language-to-SQL (NLSQL) performance.
Key Responsibilities
1. Relational Data Modeling
• Design realistic, normalized relational data models (dimension and fact tables) for four
industry verticals: Financial, Healthcare, Manufacturing, and Retail.
• Ensure schema design supports representative industry use cases and natural-language
querying patterns without exposing PHI/PII or sensitive data.
• Map data requirements to align with the Oracle AI Database Agent metadata indexing
and query execution engine.
2. Synthetic Data Generation & Loading
• Develop and execute reusable Python/SQL scripts to programmatically synthesize data
at scale.
• Generate target volume thresholds (50-500 rows per dimension table;
10,000-100,000 rows per fact table per industry).
• Perform one-time seed data loads into designated schemas within the shared Oracle
Autonomous Database.
3. Query Validation & Golden Dataset Development
• Validate data quality, primary/foreign key integrity, and representative query coverage
across all four schemas.
• Collaborate with the AI/Gemini engineering team to build prompt libraries, golden
datasets, and sample Critical User Journeys (CUJs).
• Perform end-to-end testing (Question SQL Result Execution) to confirm the
accuracy of generated SQL queries.
• Participate in cross-industry isolation testing to ensure dataset security across schema
boundaries.
4. Post-Deployment & UAT Support
• Assist with User Acceptance Testing (UAT) bug fixing and schema refinements based on
demo feedback.