Job Summary
The Data Scientist will analyze large, complex operational datasets to develop statistical insights, machine learning solutions, and executive-ready dashboards that support business and product decisions. The role requires strong proficiency in Python, SQL Server, Oracle, Tableau, Splunk, statistical inference, and machine learning fundamentals. The candidate will translate ambiguous business requirements into actionable data problems, build production-quality data pipelines and dashboards, monitor deployed solutions for performance issues and data drift, and communicate findings effectively to Product Owners, VPs, and other stakeholders.
Key Responsibilities
• Analyze large and complex operational datasets using Python, SQL, and statistical methods to identify trends, patterns, and actionable insights.
• Develop data analysis and machine learning solutions using Python libraries such as pandas, NumPy, scikit-learn, and at least one of statsmodels, TensorFlow, or PyTorch.
• Apply statistical inference and hypothesis testing to investigate business and operational questions.
• Develop and evaluate machine learning models for classification, regression, clustering, and tree-based modeling.
• Query and analyze data from SQL Server and Oracle databases, with exposure to Snowflake where applicable.
• Build, optimize, and maintain Tableau dashboards that are user-friendly, performant, and suitable for executive audiences.
• Develop Splunk SPL searches and dashboards using operational data.
• Build and maintain production-quality data pipelines, dashboards, and reusable code rather than one-off scripts.
• Use Git/GitLab for version control and maintain well-documented, maintainable code.
• Monitor deployed pipelines, dashboards, and models to identify data drift, performance issues, and operational problems.
• Translate ambiguous business requests into clearly defined, actionable data problems and deliver practical solutions independently.
• Present findings, recommendations, and technical insights to non-technical stakeholders, Product Owners, and VP-level audiences.
Required Qualifications
• Strong proficiency in Python, including pandas, NumPy, and scikit-learn, plus experience with at least one of statsmodels, TensorFlow, or PyTorch.
• Experience working with SQL Server and Oracle; Snowflake experience is an advantage.
• Experience developing and optimizing Tableau dashboards for usability, performance, and executive reporting.
• Experience with Splunk SPL searches and dashboard development using operational data.
• Strong understanding of statistical inference and hypothesis testing on large, complex operational datasets.
• Knowledge of machine learning fundamentals, including classification, regression, clustering, and tree-based models.
• Experience with Git/GitLab version control and writing maintainable, documented code.
• Strong communication skills and the ability to present findings to technical and non-technical stakeholders, including senior leadership.
• Ability to work independently, clarify ambiguous business requirements, and deliver production-quality data solutions.
Preferred Qualifications
• Experience in telecommunications, cable, or call center operations.
• Knowledge of natural language processing (NLP), including transformers, Hugging Face, topic modeling, embeddings, or semantic search.
• Experience with MLOps tools and practices, including Prefect, Airflow, Docker, or GitLab CI/CD.
• Experience with time-series analysis involving high-volume event data.
• Familiarity with vector databases such as ChromaDB, Pinecone, or FAISS.
• Experience with Alteryx, including migration of workflows to Python, or tools such as Streamlit.
• Experience with REST APIs for platforms such as Jira, Tableau, or Splunk.
• Familiarity with cloud data platforms such as Google Cloud Storage (GCS), AWS, or Azure.
• Experience with Snowflake.